Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Dec 199219922001200920182026
48 results for value function learning

Improved RL value function approximation using graph-based feature learning.

problem Accurate value function approximation in high-dimensional state or action spaces.
method Representation policy iteration (RPI) with graph-based feature learning algorithms.
result Node2vec and Variational Graph Auto-Encoder outperform PVFs in low-dimensional feature space.

Geometric approach improves reinforcement learning representation.

problem Improving reinforcement learning representation learning.
method Formal evidence through geometric properties of value functions.
result Optimizing value functions reduces to predicting adversarial value functions (AVFs).

POLO framework enables efficient learning and exploration in model-based control.

problem Efficient learning and exploration in model-based control settings.
method Combines local model-based control, global value function learning, and exploration.
result POLO framework accelerates value function learning and enables better policies.

Develops hierarchical reinforcement learning value function approximators.

problem Estimating long-term returns in reinforcement learning with multiple goals.
method Introduces hierarchical universal value function approximators (H-UVFAs) using the options framework.
result Demonstrates generalization and improved performance of H-UVFAs over UVFAs.

Paper presents an efficient exploration method for reinforcement learning.

problem Efficient exploration in reinforcement learning with uncertainty quantification.
method Parameterized Indexed Value Function (PIV) using index sampling.
result Proves the regret bound for learning PIV in a tabular setting and proposes PINs for computational learning.

UA-LQE improves value function learning by selectively erasing uncertain entries in Q-matrix.

problem Improving value function learning in complex reinforcement learning tasks.
method Uncertainty-aware low-rank Q-matrix estimation (UA-LQE) algorithm.
result UA-LQE selectively erases uncertain entries in Q-matrix to improve value function approximation.

We study the use of randomized value functions to guide deep exploration in reinforcement learning. This offers an elegant means for synthesizing statistically and computationally efficient exploration with common practical approaches to value function learning. We present several reinforcement learning algorithms that…

2017-03-22abs ↗pdf ↗

PBVFs generalize across policies using learned value functions.

problem RL algorithms forget information about old policies when updating value functions to track the learned policy.
method Introduce Parameter-Based Value Functions (PBVFs) that include policy parameters in their inputs, enabling them to generalize across different policies.
result PBVFs enable zero-shot learning of new policies that outperform any policy seen during training.

Estimating the value function for a fixed policy is a fundamental problem in reinforcement learning. Policy evaluation algorithms---to estimate value functions---continue to be developed, to improve convergence rates, improve stability and handle variability, particularly for off-policy learning. To understand the prop…

2018-08-28abs ↗pdf ↗

New RL method learns value function for many policies using few key states.

problem Evaluate and improve policies in continuous control problems.
method Combines actor-critic architecture and policy embedding to learn a single value function for many policies.
result Value function minimizes prediction error by learning a small set of 'probing states' and their impact on policies' returns.

This paper explores a new DRL algorithm that approximates both state-value and state-action functions.

problem Overestimation bias in state-action value function.
method Developed and analyzed the Deep Quality-Value (DQV) algorithm to approximate both VV and QQ functions.
result DQV and DQV-Max algorithms perform better due to less overestimation bias in QQ function.

Paper proposes a new adaptive multiscale value function approximation for reinforcement learning.

problem Value function approximation in reinforcement learning with varying complexity.
method Adaptive multiscale approximation using multiresolution analysis and tree approximation.
result Convergence rate of the multiscale approximation is independent of basis function regularity.

Symbolic regression constructs smooth value functions for reinforcement learning.

problem Function approximators in reinforcement learning are black-box models with hyper-parameter tuning.
method Symbolic regression methods for constructing smooth value functions in the form of analytic expressions.
result Symbolic regression methods yield well-performing policies and are compact and mathematically tractable.

Novel framework for Bayesian reinforcement learning infers value function distributions.

problem Bayesian reinforcement learning's challenges in inferring value function distributions.
method Inferential Induction framework for Bayesian reinforcement learning, developing Bayesian Backwards Induction algorithm.
result Proposed algorithm is competitive with state-of-the-art methods.

A new method separates long-term value functions into components for better reinforcement learning.

problem Learning long-term goals in reinforcement learning settings with temporal discounting.
method TD(ΔΔ) learning, which breaks down value functions into components based on discount factors.
result TD(ΔΔ) learning improves scalability and performance over standard TD learning in certain settings.

Paper introduces a new value function for state transitions and optimal policy learning.

problem Learning optimal policies from state transitions and actions.
method Develops a forward dynamics model to maximize a novel value function Q(s,s)Q(s, s').
result Demonstrates benefits in value function transfer, redundant action spaces, and off-policy learning.

Self-consistent models improve reinforcement learning by aligning predictions with future values.

problem Improving reinforcement learning by aligning model predictions with future values.
method Proposes multiple self-consistency updates to encourage a learned model and value function to be consistent with each other.
result Self-consistency helps both policy evaluation and control in both tabular and function approximation settings.

We present a framework to derive risk bounds for vector-valued learning with a broad class of feature maps and loss functions. Multi-task learning and one-vs-all multi-category learning are treated as examples. We discuss in detail vector-valued functions with one hidden layer, and demonstrate that the conditions under…

2016-06-05abs ↗pdf ↗

A new method for estimating joint value functions in multi-scene reinforcement learning.

problem High variance in samples for policy gradient computations in multi-scene environments.
method Sparse attention mechanism over multiple value function hypotheses to approximate the true joint value function.
result Significant improvements in reward scores and enhanced navigation efficiency across OpenAI ProcGen environments.

New kernels capture both local and non-local interactions efficiently.

problem Designing kernels that capture both local and non-local interactions while remaining computationally tractable.
method Spectral truncation kernels based on CC^*-algebra.
result Spectral truncation kernels induce interactions across the data function domain and reduce computational cost.

This paper interpolates reward functions to predict optimal value functions in MORL.

problem Finding optimal value functions in MORL requires recomputing for each set of weights.
method Interpolating reward function weights to smooth value function transformations.
result Smooth interpolation of optimal value functions over reward function weights.

This paper introduces a new framework for learning nearly decomposable Q-functions via communication minimization.

problem Challenges in multi-agent reinforcement learning, especially scalability and non-stationarity.
method Learning nearly decomposable Q-functions (NDQ) via communication minimization, introducing two information-theoretic regularizers.
result Significantly outperforms baseline methods on the StarCraft unit micromanagement benchmark.

Paper develops efficient RL algorithm for general value function approximation.

problem Lack of theory for RL with general value function approximation.
method Provable efficient RL algorithm using bounded eluder dimension.
result Achieves a regret bound of O~(poly(dH)T)\widetilde{O}(\mathrm{poly}(dH)\sqrt{T}).

A reinforcement learning framework combining value function and tree search planner for strategic and tactical decisions.

problem Strategic and tactical decision-making in discrete environments.
method Combines value function and tree search planner, using uncertainty modeling and risk measurement.
result Improves performance and learning speed on hard exploration environments.

The paper explores when and why value decomposition algorithms work in cooperative multi-agent reinforcement learning.

problem The applicability and convergence properties of value decomposition algorithms in cooperative multi-agent reinforcement learning are unclear.
method The paper introduces decomposable games and proves that applying the multi-agent fitted Q-Iteration algorithm leads to an optimal Q-function in these games.
result The paper offers theoretical insights into when and why value decomposition algorithms converge in cooperative multi-agent reinforcement learning.

New algorithm learns value and advantage functions for continuous-time Markov processes without structural assumptions.

problem Learning value and advantage functions for continuous-time Markov processes without structural assumptions.
method Proposes Sobolev-prox fitted qq-learning algorithm based on Hilbert-space positive definiteness and boundedness properties of Bellman operators.
result Identifies ellipticity as a key structural property enabling reinforcement learning for Markov diffusions.

We propose randomized least-squares value iteration (RLSVI) -- a new reinforcement learning algorithm designed to explore and generalize efficiently via linearly parameterized value functions. We explain why versions of least-squares value iteration that use Boltzmann or epsilon-greedy exploration can be highly ineffic…

2014-02-04abs ↗pdf ↗

The paper develops methods to estimate value functions in reinforcement learning for infinite horizon problems.

problem Constructing confidence intervals for value functions in infinite horizon reinforcement learning.
method Modeling the Q-function using series/sieve method and recursively updating the policy using SAVE method.
result The proposed CI achieves nominal coverage even when the optimal policy is not unique.

Positive definite operator-valued kernels generalize the well-known notion of reproducing kernels, and are naturally adapted to multi-output learning situations. This paper addresses the problem of learning a finite linear combination of infinite-dimensional operator-valued kernels which are suitable for extending func…

2012-03-07abs ↗pdf ↗

UVU simplifies value uncertainty quantification in RL.

problem Estimating epistemic uncertainty in value functions for reinforcement learning.
method UVU uses squared prediction errors between an online learner and a fixed, randomly initialized target network, incorporating policy-conditional value uncertainty.
result UVU achieves equal performance to large ensembles on challenging offline RL settings, with computational savings.

The paper proposes a method to decompose value functions in RL for better understanding and prediction.

problem Understanding and predicting the dynamics and returns in reinforcement learning models.
method A two-step approach decomposing the value function into future dynamics and trajectory returns, with a practical deep RL algorithm.
result The proposed algorithm outperforms in MuJoCo tasks, especially under delayed reward settings.

A novel approach uses an ensemble of Gaussian processes for robust and adaptive reinforcement learning.

problem Adaptive reinforcement learning in large or continuous state spaces.
method Online scalable (OS) approach with a weighted ensemble of Gaussian processes.
result The ensemble approach improves performance in adversarial settings.

In this paper we consider the problems of supervised classification and regression in the case where attributes and labels are functions: a data is represented by a set of functions, and the label is also a function. We focus on the use of reproducing kernel Hilbert space theory to learn from such functional data. Basi…

2015-10-28abs ↗pdf ↗

PPG separates policy and value function training phases for better reinforcement learning efficiency.

problem Challenges in traditional reinforcement learning methods for policy and value function optimization.
method Integrates Phasic Policy Gradient framework that splits policy and value function training into distinct phases.
result Significantly improves sample efficiency on Procgen Benchmark compared to PPO.