SDM Policy accelerates inference for robotic tasks while maintaining high action quality.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Holographic principle matches deformed Liouville theory action.
QAM uses adjoint matching to optimize continuous-action RL policies efficiently.
New methods use vector search and nearest-neighbor matching for policy learning in causal inference.
New method AM learns optimal vector fields for entire distribution sequences, matching OT.
We consider the problem of learning the optimal action-value function in the discounted-reward Markov decision processes (MDPs). We prove a new PAC bound on the sample-complexity of model-based value iteration algorithm in the presence of the generative model, which indicates that for an MDP with N state-action pairs a…
We present a complete acyclic matching of the Hasse diagram associated with the face lattice of a hypersimplex. Since a hypersimplex is a convex polytope, there is a natural way to form a CW complex from its faces. We will then utilize this matching along with discrete Morse theory and some topological techniques to cl…
Poisson actions of Poisson Lie groups have an interesting and rich geometric structure. We will generalize some of this structure to Dirac actions of Dirac Lie groups. Among other things, we extend a result of Jiang-Hua-Lu, which states that the cotangent Lie algebroid and the action algebroid for a Poisson action form…
A new framework for offline RL improves policy flexibility and regularity.
Paper computes optimal matching between curves on manifolds.
Identifies LA-groups via VB-group structure and complementary actions.
We study the local structure of Lie bialgebroids at regular points. In particular, we classify all transitive Lie bialgebroids. In special cases, they are connected to classical dynamical -matrices and matched pairs induced by Poisson group actions
A new calibration metric bridges testability and actionability.
Unified framework for imitation learning via moment matching.
Bayesian method learns optimal momentum for landmark matching.
Diffusion models mimic human actions in sequential tasks.
Noiseless IO bounds inferred from demonstrations, matching adversarial settings.
A scheme robust to action erasures improves MAB performance.
Study modular class of Lie ∞-algebroids and their adjoint actions.
This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of known action features confounded by an non-linear action-independent term. We design new algorithms that achieve regret …
We consider an adversarial online learning setting where a decision maker can choose an action in every stage of the game. In addition to observing the reward of the chosen action, the decision maker gets side observations on the reward he would have obtained had he chosen some of the other actions. The observation str…
Paper introduces Decentralized Non-stationary Competing Bandits ( exttt{DNCB}) for dynamic matching markets.
In urban environments, supply resources have to be constantly matched to the "right" locations (where customer demand is present) so as to improve quality of life. For instance, ambulances have to be matched to base stations regularly so as to reduce response time for emergency incidents in EMS (Emergency Management Sy…
Improved GFlowNets learn more efficiently with trajectory balance.
A quantum field theory for Spin(7)-instantons derived from moduli spaces.
Study non-standard bi-orders on punctured torus bundles, matching standard ones in key subgroups.
This paper provides theoretical foundations for using quantized actions in behavior cloning.
Thompson sampling, a Bayesian method for balancing exploration and exploitation in bandit problems, has theoretical guarantees and exhibits strong empirical performance in many domains. Traditional Thompson sampling, however, assumes perfect compliance, where an agent's chosen action is treated as the implemented actio…
Researchers use shape analysis to recover protein structures from Cryo-EM data.
We propose Generative Predecessor Models for Imitation Learning (GPRIL), a novel imitation learning algorithm that matches the state-action distribution to the distribution observed in expert demonstrations, using generative models to reason probabilistically about alternative histories of demonstrated states. We show …
We study online reinforcement learning for finite-horizon deterministic control systems with {\it arbitrary} state and action spaces. Suppose that the transition dynamics and reward function is unknown, but the state and action space is endowed with a metric that characterizes the proximity between different states and…
We provide a simple and efficient algorithm for adversarial -action -outcome non-degenerate locally observable partial monitoring game for which the -round minimax regret is bounded by , matching the best known information-theoretic upper bound. The same algorithm also achieves…
We consider a stochastic multi-armed bandit (MAB) problem with delayed impact of actions. In our setting, actions taken in the past impact the arm rewards in the subsequent future. This delayed impact of actions is prevalent in the real world. For example, the capability to pay back a loan for people in a certain socia…
New framework for reinforcement learning with sporadic state observations.
Policy gradient methods have enjoyed great success in deep reinforcement learning but suffer from high variance of gradient estimates. The high variance problem is particularly exasperated in problems with long horizons or high-dimensional action spaces. To mitigate this issue, we derive a bias-free action-dependent ba…
This paper presents a multi-staged approach to nonmyopic adaptive Gaussian process optimization (GPO) for Bayesian optimization (BO) of unknown, highly complex objective functions that, in contrast to existing nonmyopic adaptive BO algorithms, exploits the notion of macro-actions for scaling up to a further lookahead t…
The paper extends LDDMM framework to include Lie group actions in large deformation shape registration.
We address reinforcement learning problems with finite state and action spaces where the underlying MDP has some known structure that could be potentially exploited to minimize the exploration rates of suboptimal (state, action) pairs. For any arbitrary structure, we derive problem-specific regret lower bounds satisfie…
MobILE learns from expert demonstrations without access to actions, achieving strong performance guarantees.
New categorical actions link topological and algebraic structures.
We propose a framework for modeling and estimating the state of controlled dynamical systems, where an agent can affect the system through actions and receives partial observations. Based on this framework, we propose the Predictive State Representation with Random Fourier Features (RFFPSR). A key property in RFF-PSRs …
Improved exploration in cooperative multi-agent reinforcement learning.
Many problems at the intersection of combinatorics and computer science require solving for a permutation that optimally matches, ranks, or sorts some data. These problems usually have a task-specific, often non-differentiable objective function that data-driven algorithms can use as a learning signal. In this paper, w…
Proposes a new method to estimate continuous treatment policies and match treatments effectively.
New algorithm reduces switching costs in multinomial logit bandit problems.
We introduce Mix&Match (M&M) - a training framework designed to facilitate rapid and effective learning in RL agents, especially those that would be too slow or too challenging to train otherwise. The key innovation is a procedure that allows us to automatically form a curriculum over agents. Through such a curriculum …
We introduce the factored bandits model, which is a framework for learning with limited (bandit) feedback, where actions can be decomposed into a Cartesian product of atomic actions. Factored bandits incorporate rank-1 bandits as a special case, but significantly relax the assumptions on the form of the reward function…
Building agents to interact with the web would allow for significant improvements in knowledge understanding and representation learning. However, web navigation tasks are difficult for current deep reinforcement learning (RL) models due to the large discrete action space and the varying number of actions between the s…