Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4208391,2591,678 · Jun 202019922001200920182026
48 results for approximate reinforcement learning

Paper proposes a new adaptive multiscale value function approximation for reinforcement learning.

problem Value function approximation in reinforcement learning with varying complexity.
method Adaptive multiscale approximation using multiresolution analysis and tree approximation.
result Convergence rate of the multiscale approximation is independent of basis function regularity.

Efficient offline reinforcement learning with neural networks using differentiable function approximation.

problem Statistical efficiency of offline reinforcement learning with function approximators.
method Pessimistic fitted Q-learning (PFQL) and differentiable function approximation.
result Provably efficient offline reinforcement learning with differentiable function approximation.

Develops hierarchical reinforcement learning value function approximators.

problem Estimating long-term returns in reinforcement learning with multiple goals.
method Introduces hierarchical universal value function approximators (H-UVFAs) using the options framework.
result Demonstrates generalization and improved performance of H-UVFAs over UVFAs.

New findings show good representations alone are insufficient for efficient reinforcement learning.

problem Understanding when good representations are enough for efficient reinforcement learning.
method Statistical analysis of reinforcement learning methods, focusing on value-based, model-based, and policy-based learning.
result Hard thresholds for reinforcement learning methods show good representations alone are insufficient, unless they meet certain quality criteria.

Bayesian meta-reinforcement learning improves over point estimates with Laplace approximation.

problem Improving meta-reinforcement learning by providing full posterior distributions.
method Augmenting point estimates with Laplace approximation for full posterior distributions.
result Our method performs similarly to variational baselines with fewer parameters.

The paper emphasizes the importance of value distribution in reinforcement learning.

problem The focus on value expectation in reinforcement learning is insufficient.
method Developed a new algorithm based on the distributional perspective of reinforcement learning.
result Demonstrated significant distributional instability in reinforcement learning control.

Unified approach to dynamic programming improves reinforcement learning performance.

problem Improving approximate dynamic programming for broader reinforcement learning applications.
method Proposes Generalized Value Iteration (GVI) and its approximated version, Approximate GVI (AGVI), unifying value iteration, advantage learning, and dynamic policy programming.
result Demonstrates performance guarantees for AGVI, including those for existing algorithms.

Paper robustifies reinforcement learning with risk-averse methods.

problem Making predictions robust to changes in system dynamics or rewards.
method Approximates Robust Reinforcement Learning using ΦΦ-divergence and Risk-Averse formulation.
result Classical Reinforcement Learning can be robustified using standard deviation penalization.

The paper tackles hierarchical reinforcement learning by approximating optimal solutions for the Traveling Salesman Problem.

problem Approximating optimal solutions for the Traveling Salesman Problem using hierarchical reinforcement learning.
method Mapping the problem into a Reward Discounted Traveling Salesman Problem and deriving approximate solutions using local policies.
result Three stochastic policies are proposed that guarantee better performance than any deterministic policy.

This work analyzes nonexpansive stochastic approximations with Markovian noise, proving convergence in reinforcement learning.

problem Applying stochastic approximation to reinforcement learning settings with nonexpansive operators.
method Investigates nonexpansive stochastic approximations with Markovian noise, providing asymptotic and finite sample analysis.
result First-time proof of convergence for classical tabular average reward temporal difference learning.

New Zap Q-learning accelerates reinforcement learning with neural networks.

problem Accelerate convergence of reinforcement learning algorithms.
method Introduces a new framework for analysis of stochastic approximation algorithms, proving consistency under non-degeneracy assumption.
result Zap Q-learning with neural network function approximation converges quickly and is robust to function approximation architecture choice.

Approximate models help RL by reducing policy search space.

problem How much does an approximate model help in learning near-optimal policies in RL?
method Study sample complexity in RL with an approximate model, providing an algorithm and a lower bound.
result An approximate model can reduce sample complexity by eliminating sub-optimal actions.

The paper analyzes convergence rates for stochastic approximation and reinforcement learning.

problem Establishing almost sure convergence rates for stochastic approximation and reinforcement learning under Markovian noise.
method A novel Lyapunov drift construction that applies a Poisson-equation based correction for Markovian noise to the Moreau-envelope smoothing for contractive mappings.
result Almost sure convergence rates for specific learning rates are derived, with rates arbitrarily close to o(n12η)o(n^{1 - 2η}) and o(n1)o(n^{-1}).

Enhances RL with function approximation, improving regret bounds.

problem Improving exploration in reinforcement learning with function approximation.
method Prior-dependent Bayesian regret bound for PSRL with linear mixture MDPs, using value-targeted model learning and variance reduction.
result Established an upper bound of O(dH3TlogT){\mathcal{O}}(d\sqrt{H^3 T \log T}) for PSRL.

Paper establishes convergence rates and concentration bounds for stochastic approximation and reinforcement learning with Markovian noise.

problem Analyzing convergence rates and concentration bounds for stochastic approximation and reinforcement learning with Markovian noise.
method Novel discretization of the mean ODE of stochastic approximation algorithms using intervals with diminishing length.
result First almost sure convergence rate and maximal concentration bound with exponential tails for contractive stochastic approximation algorithms with Markovian noise.

Improved RL value function approximation using graph-based feature learning.

problem Accurate value function approximation in high-dimensional state or action spaces.
method Representation policy iteration (RPI) with graph-based feature learning algorithms.
result Node2vec and Variational Graph Auto-Encoder outperform PVFs in low-dimensional feature space.

New algorithms predict reinforcement learning values efficiently.

problem Predicting reinforcement learning values with linear function approximation.
method Multi-timescale stochastic approximation of cross entropy method.
result Proved convergence and achieved good performance in experiments.

The paper tackles cooperative RL with function approximation, achieving near-optimal learning with limited communication.

problem Cooperative multi-agent reinforcement learning with function approximation.
method Careful message-passing and cooperative value iteration.
result Achieving near-optimal no-regret learning with limited communication in cooperative multi-agent settings.

This paper compares expected and distributional reinforcement learning methods.

problem Understanding why distributional reinforcement learning performs better than expected reinforcement learning.
method Analyzes differences in tabular, linear, and non-linear approximation settings.
result Distributional RL can hurt performance if it does not induce identical behavior.

New algorithm learns value and advantage functions for continuous-time Markov processes without structural assumptions.

problem Learning value and advantage functions for continuous-time Markov processes without structural assumptions.
method Proposes Sobolev-prox fitted qq-learning algorithm based on Hilbert-space positive definiteness and boundedness properties of Bellman operators.
result Identifies ellipticity as a key structural property enabling reinforcement learning for Markov diffusions.

Unified algorithm for reinforcement learning with function approximation.

problem Limited scalability of Q(σ,λ) for large-scale learning.
method Proposes GQ(σ,λ) with linear function approximation to extend tabular Q(σ,λ).
result Empirical results show GQ(σ,λ) outperforms full-sampling and pure-expectation methods.

Improves imitation learning in RL by learning reward function efficiently.

problem Lack of effective reward function approximation in AIRL for imitation tasks.
method Proposes Off-Policy AIRL that combines adversarial learning with efficient reward function approximation.
result Shows superior imitation performance and efficiency compared to state-of-the-art AIL algorithms.

Paper tackles reinforcement learning for STL specifications with state history.

problem Learning optimal policies to satisfy STL specifications often requires too much state history, making the problem computationally intractable.
method Proposes a compact augmented state-space representation to capture state history and an approximation method to solve the objective.
result Shows the performance bound of the approximate solution and compares it with an existing technique.

Safe reinforcement learning with nonconvex constraints using convex approximations.

problem Safe reinforcement learning with nonlinear function approximation.
method Constructing surrogate convex constrained optimization problems by replacing nonconvex functions with convex quadratic functions.
result Solutions to surrogate problems converge to a stationary point of the original nonconvex problem.

We study reinforcement learning under model misspecification, where we do not have access to the true environment but only to a reasonably close approximation to it. We address this problem by extending the framework of robust MDPs to the model-free Reinforcement Learning setting, where we do not have access to the mod…

2017-06-15abs ↗pdf ↗

New algorithm FLUTE achieves uniform-PAC convergence in RL with linear approx.

problem RL with linear function approximation lacks uniform-PAC guarantees.
method FLUTE algorithm with minimax value function estimator and multi-level partition scheme.
result Uniform-PAC convergence to optimal policy with high probability.

Kernel-based function approximation improves reinforcement learning performance.

problem Average reward reinforcement learning in infinite horizon settings.
method Optimistic algorithm based on kernel ridge regression.
result No-regret performance guarantees and confidence intervals for kernel-based predictions.

Paper analyzes distributional reinforcement learning with value function approximation, introducing Bellman unbiasedness and a new algorithm.

problem Improving reinforcement learning by capturing environmental stochasticity and addressing infinite dimensionality.
method Introduces Bellman unbiasedness and proposes SF-LSVI algorithm for provably efficient distributional reinforcement learning.
result Achieves a tight regret bound of O(d_E H^3/2 √K) for distributional reinforcement learning.

Unified view of federated learning and distributed RL using local stochastic approximation.

problem Finding the root of an operator composed of local operators in a network of agents with dependent data.
method Local stochastic approximation over a network of agents with Markov process-dependent data.
result Convergence rates of local stochastic approximation for both constant and time-varying step sizes, within a logarithmic factor of independent data.

New algorithm combines Cramér distance with function approximation for distributional reinforcement learning.

problem Limited theoretical understanding of practical distributional reinforcement learning methods.
method Adapts Cramér distance to arbitrary vectors, derives new distributional algorithm combining Cramér-based and function approximation.
result First proof of convergence for a distributional algorithm combined with function approximation.

State2vec improves RL by learning state embeddings that generalize across policies.

problem Inefficient generalization across policies in RL.
method Extends node2vec to learn state embeddings accounting for discounted future state transitions.
result Captures the geometry of the state space, leading to sample-efficient value function approximation.

Investigates Q-learning bottlenecks with function approximation and sampling methods.

problem Understanding and mitigating issues in Q-learning with function approximation.
method Unit testing framework with oracles to disentangle sources of error; novel sampling method based on function approximation error.
result Large neural networks improve learning stability and offer practical compensations for overfitting.

A novel Q-learning variant reduces underestimation bias in deep actor-critic methods for reinforcement learning.

problem Underestimation bias in deep actor-critic methods for reinforcement learning.
method Introduces a parameter-free Q-learning variant that combines maximum and minimum operators to bound value estimates.
result Improves state-of-the-art performance on OpenAI Gym tasks.

Novel framework for Bayesian reinforcement learning infers value function distributions.

problem Bayesian reinforcement learning's challenges in inferring value function distributions.
method Inferential Induction framework for Bayesian reinforcement learning, developing Bayesian Backwards Induction algorithm.
result Proposed algorithm is competitive with state-of-the-art methods.

New algorithms improve distributional TD learning with linear approximations.

problem Estimating return distributions in reinforcement learning.
method Fine-grained analysis of linear-categorical Bellman equation, variance reduction techniques.
result Tight sample complexity bounds for distributional TD learning with linear approximations.

Improved reinforcement learning algorithm with linear approximation for unknown dynamics.

problem Reinforcement learning with adversarial changing cost functions and bandit feedback.
method Combines mirror-descent and least squares policy evaluation in an auxiliary MDP.
result Obtains an O~(K6/7)\widetilde O(K^{6/7}) regret bound, significantly improving over previous methods.

FFN addresses spectral bias in neural value approximation, improving reinforcement learning performance.

problem Spectral bias in neural value approximation, leading to slow convergence and poor performance.
method Proposes Fourier feature networks (FFN) to overcome spectral bias by using a composite neural tangent kernel.
result FFN achieves state-of-the-art performance on challenging continuous control domains with faster convergence and better stability.

Transformers achieve near-optimal dynamic regret in non-stationary reinforcement learning.

problem Understanding and handling non-stationary environments in reinforcement learning.
method Demonstrated that transformers can achieve nearly optimal dynamic regret bounds in non-stationary settings.
result Transformers can approximate and learn strategies for non-stationary environments, matching or outperforming existing expert algorithms.

New algorithms improve reinforcement learning with multi-step greedy policies.

problem Difficulty in monotonic policy improvement with soft-policy updates.
method Formulated and analyzed online and approximate algorithms using multi-step greedy operators.
result Guaranteed monotonic policy improvement with sufficiently large update stepsize.