Optimal estimator derived for partially observable LTI systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian approach optimizes in-context learning for state space models.
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
New RL method handles large state-action spaces with complex models.
Optimal transport theory applied to quantum states on Grassmannians.
This paper deals with discrete-time Markov control processes on a general state space. A long-run risk-sensitive average cost criterion is used as a performance measure. The one-step cost function is nonnegative and possibly unbounded. Using the vanishing discount factor approach, the optimality inequality and an optim…
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
We propose an algorithm for deterministic continuous Markov Decision Processes with sparse rewards that computes the optimal policy exactly with no dependency on the size of the state space. The algorithm has time complexity of and memory complexity of , where is the…
Generically learns movement control policies from exploration data.
Efficient RL in large POMDPs with latent determinism and embeddings.
We present a novel technique to solve the problem of managing optimally a pumped hydroelectric storage system. This technique relies on representing the system as a stochastic optimal control problem with state constraints, these latter corresponding to the finite volume of the reservoirs. Following the recent level-se…
Dynamical systems with large state-spaces are often expensive to thoroughly explore experimentally. Coarse-graining methods aim to define simpler systems which are more amenable to analysis and exploration; most current methods, however, focus on a priori state aggregation based on similarities in transition rates, whi…
Paper solves POMDPs in continuous time and discrete spaces.
Unified framework extends adjoint Schrödinger bridge sampler to discrete spaces.
A DRL framework optimizes portfolios using a LFSS module for feature extraction.
AE-LSVI identifies near-optimal policies in complex systems with minimal data.
Sample inefficiency is a long-lasting problem in reinforcement learning (RL). The state-of-the-art estimates the optimal action values while it usually involves an extensive search over the state-action space and unstable optimization. Towards the sample-efficient RL, we propose ranking policy gradient (RPG), a policy …
The paper tackles reinforcement learning with exogenous variables and rewards.
Operator calculus for population-based optimization provides a unified framework for analyzing convergence of various methods.
Active learning selects inputs for GPSSM to learn latent states.
We study online reinforcement learning for finite-horizon deterministic control systems with {\it arbitrary} state and action spaces. Suppose that the transition dynamics and reward function is unknown, but the state and action space is endowed with a metric that characterizes the proximity between different states and…
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
Paper introduces multitask neural networks for efficient stochastic control problems.
Develops a regression approach for solving MDPs with general state and action spaces.
New RL method reduces sample complexity for large state-action spaces.
New method learns state embeddings from demonstrations for improved reinforcement learning.
SMEs provide a transparent testbed for RL evaluation.
This work tackles large action spaces in RL by binarizing actions.
Predictability enables efficient parallelization of nonlinear models.
A new method shapes reinforcement learning environments by abstracting large state spaces.
Using stochastic gradient search and the optimal filter derivative, it is possible to perform recursive (i.e., online) maximum likelihood estimation in a non-linear state-space model. As the optimal filter and its derivative are analytically intractable for such a model, they need to be approximated numerically. In [Po…
A nonparametric approach for policy learning for POMDPs is proposed. The approach represents distributions over the states, observations, and actions as embeddings in feature spaces, which are reproducing kernel Hilbert spaces. Distributions over states given the observations are obtained by applying the kernel Bayes' …
We present the first PAC optimal algorithm for Bayes-Adaptive Markov Decision Processes (BAMDPs) in continuous state and action spaces, to the best of our knowledge. The BAMDP framework elegantly addresses model uncertainty by incorporating Bayesian belief updates into long-term expected return. However, computing an e…
A method for robust reinforcement learning in large state spaces.
The paper tackles finding optimal treatment sequences in continuous state spaces.
We present an efficient algorithm for model-free episodic reinforcement learning on large (potentially continuous) state-action spaces. Our algorithm is based on a novel -learning policy with adaptive data-driven discretization. The central idea is to maintain a finer partition of the state-action space in regions w…
Hybrid model improves sequential data prediction by combining neural and time series models.
Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model. We propose a parametric Q-learning algorithm that finds an approximate-optimal policy using a sample size proportional to the feature dimension and invariant …
A new active learning method for Gaussian process models.
New optimizer G-AdaGrad improves upon AdaGrad for non-convex machine learning problems.
We develop a normative framework for hierarchical model-based policy optimization based on applying second-order methods in the space of all possible state-action paths. The resulting natural path gradient performs policy updates in a manner which is sensitive to the long-range correlational structure of the induced st…
New RL method reduces sample complexity for large policy spaces.
RPO uses past and future state-action info for better policy optimization.
Develops a learning model predictive controller for competitive racing.
Most real-world problems have huge state and/or action spaces. Therefore, a naive application of existing tabular solution methods is not tractable on such problems. Nonetheless, these solution methods are quite useful if an agent has access to a relatively small state-action space homomorphism of the true environment …
In this paper, we introduce a novel form of value function, , that expresses the utility of transitioning from a state to a neighboring state and then acting optimally thereafter. In order to derive an optimal policy, we develop a forward dynamics model that learns to make next-state predictions that…
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
State-space models are used in a wide range of time series analysis formulations. Kalman filtering and smoothing are work-horse algorithms in these settings. While classic algorithms assume Gaussian errors to simplify estimation, recent advances use a broader range of optimization formulations to allow outlier-robust e…