The paper explores learning good policies from past data in large state spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficient RL in large POMDPs with latent determinism and embeddings.
New RL method reduces sample complexity for large state-action spaces.
A method for robust reinforcement learning in large state spaces.
This paper simplifies OPE in large state spaces using state abstractions.
New RL method handles large state-action spaces with complex models.
This work tackles large action spaces in RL by binarizing actions.
A new method shapes reinforcement learning environments by abstracting large state spaces.
We consider Markov Decision Processes (MDPs) where the rewards are unknown and may change in an adversarial manner. We provide an algorithm that achieves state-of-the-art regret bound of , where is the state space, is the action space, is the mixing time of the MDP, and $…
Reinforcement learning (RL) in Markov decision processes (MDPs) with large state spaces is a challenging problem. The performance of standard RL algorithms degrades drastically with the dimensionality of state space. However, in practice, these large MDPs typically incorporate a latent or hidden low-dimensional structu…
New RL method reduces sample complexity for large policy spaces.
We provide a comprehensive overview and tooling for GP modeling with non-Gaussian likelihoods using state space methods. The state space formulation allows for solving one-dimensional GP models in time and memory complexity. While existing literature has focused on the connection between GP regression …
We consider the problem of learning low-dimensional representations for large-scale Markov chains. We formulate the task of representation learning as that of mapping the state space of the model to a low-dimensional state space, called the kernel space. The kernel space contains a set of meta states which are desired …
ZoomRL learns efficient strategies for large state-action spaces using a metric.
A new method learns complex dynamical systems from data efficiently.
We forecast S&P 500 excess returns using a flexible Bayesian econometric state space model with non-Gaussian features at several levels. More precisely, we control for overparameterization via novel global-local shrinkage priors on the state innovation variances as well as the time-invariant part of the state space mod…
We present a scalable and robust Bayesian inference method for linear state space models. The method is applied to demand forecasting in the context of a large e-commerce platform, paying special attention to intermittent and bursty target statistics. Inference is approximated by the Newton-Raphson algorithm, reduced t…
Adaptive discretization improves model-based RL in large spaces.
In complex tasks, such as those with large combinatorial action spaces, random exploration may be too inefficient to achieve meaningful learning progress. In this work, we use a curriculum of progressively growing action spaces to accelerate learning. We assume the environment is out of our control, but that the agent …
Consider the Chern-Simons topological quantum field theory with gauge group SU(2) and level k. Given a knot in the 3-sphere, this theory associates to the knot exterior an element in a vector space. We call this vector the knot state and study its asymptotic properties when the level is large. The latter vector space b…
New RL algorithm tackles large state spaces using optimistic function approximation.
State representation learning (SRL) in partially observable Markov decision processes has been studied to learn abstract features of data useful for robot control tasks. For SRL, acquiring domain-agnostic states is essential for achieving efficient imitation learning. Without these states, imitation learning is hampere…
When using reinforcement learning (RL) algorithms to evaluate a policy it is common, given a large state space, to introduce some form of approximation architecture for the value function (VF). The exact form of this architecture can have a significant effect on the accuracy of the VF estimate, however, and determining…
Combines pseudo-point and state space approximations for scalable GPs.
Dynamical systems with large state-spaces are often expensive to thoroughly explore experimentally. Coarse-graining methods aim to define simpler systems which are more amenable to analysis and exploration; most current methods, however, focus on a priori state aggregation based on similarities in transition rates, whi…
This paper achieves first-order regret bounds in reinforcement learning with large state spaces.
We state conjectures on the asymptotic behavior of the volumes of moduli spaces of Abelian differentials and their Siegel-Veech constants as genus tends to infinity. We provide certain numerical evidence, describe recent advances and the state of the art towards proving these conjectures.
We give a proof of the celebrated stability theorem of Perelman stating that for a noncollapsing sequence of Alexandrov spaces with curvature bounded below Gromov-Hausdorff converging to a compact Alexandrov space , is homeomorphic to for all large .
A new framework using kernel packets overcomes limitations of state space models for multi-dimensional data.
New method speeds up Gaussian process inference for large datasets.
New framework improves sample efficiency and robustness in RL with smooth policies.
We consider the Markov Decision Process (MDP) of selecting a subset of items at each step, termed the Select-MDP (S-MDP). The large state and action spaces of S-MDPs make them intractable to solve with typical reinforcement learning (RL) algorithms especially when the number of items is huge. In this paper, we present …
Improved RL in BMDPs reduces regret to O(sqrt(T)+n).
In nonlinear state-space models, sequential learning about the hidden state can proceed by particle filtering when the density of the observation conditional on the state is available analytically (e.g. Gordon et al., 1993). This condition need not hold in complex environments, such as the incomplete-information equili…
Stochastic domains often involve risk-averse decision makers. While recent work has focused on how to model risk in Markov decision processes using risk measures, it has not addressed the problem of solving large risk-averse formulations. In this paper, we propose and analyze a new method for solving large risk-averse …
GUM tackles MARL by avoiding overestimation through state-marginal restriction.
Stanza models complex time series with balance between traditional and deep learning approaches.
Machine learning and quantum computing are two technologies each with the potential for altering how computation is performed to address previously untenable problems. Kernel methods for machine learning are ubiquitous for pattern recognition, with support vector machines (SVMs) being the most well-known method for cla…
New method improves Kalman filtering and smoothing for large state spaces.
Paper explores using LLMs for zero-shot reinforcement learning in continuous spaces.
Paper presents a fast method for estimating hidden states in Bayesian models.
Structured state space models improve ECG classification and reveal new insights.
A Budgeted Markov Decision Process (BMDP) is an extension of a Markov Decision Process to critical applications requiring safety constraints. It relies on a notion of risk implemented in the shape of a cost signal constrained to lie below an - adjustable - threshold. So far, BMDPs could only be solved in the case of fi…
Gromov showed that for fixed, arbitrarily large C, any uniformly C-Lipschitz affine action of a random group in his graph model on a Hilbert space has a fixed point. We announce a theorem stating that more general affine actions of the same random group on a Hilbert space have a fixed point. We discuss some aspects of …
How can we efficiently propagate uncertainty in a latent state representation with recurrent neural networks? This paper introduces stochastic recurrent neural networks which glue a deterministic recurrent neural network and a state space model together to form a stochastic and sequential neural generative model. The c…
Quantum algorithms speed up reinforcement learning policies in large state-action spaces.
We consider the problem of estimating from sample paths the absolute spectral gap of a reversible, irreducible and aperiodic Markov chain over a finite state space . We propose the (Upper Confidence Power Iteration) algorithm for this problem, a low-complexity algorithm …
We give a complete framework for the Gupta-Bleuler quantization of the free electromagnetic field on globally hyperbolic space-times. We describe one-particle structures that give rise to states satisfying the microlocal spectrum condition. The field algebras in the so-called Gupta-Bleuler representations satisfy the t…