The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method calibrates value predictions in offline RL to improve reliability.
Paper studies offline RL with linear approx, focusing on inherent Bellman error.
Polynomial-time RL algorithm for constant actions under linear Bellman completeness.
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…
New method for off-policy evaluation in POMDPs using future-dependent value functions.
This paper aims at theoretically and empirically comparing two standard optimization criteria for Reinforcement Learning: i) maximization of the mean value and ii) minimization of the Bellman residual. For that purpose, we place ourselves in the framework of policy search algorithms, that are usually designed to maximi…
In this paper, we consider the stochastic iterative counterpart of the value iteration scheme wherein only noisy and possibly biased approximations of the Bellman operator are available. We call this counterpart as the approximate value iteration (AVI) scheme. Neural networks are often used as function approximators, i…
Paper introduces v-CMC linking causality and utility.
New method quantifies uncertainty in reinforcement learning models.
Paper analyzes distributional reinforcement learning with value function approximation, introducing Bellman unbiasedness and a new algorithm.
New method stabilizes FQE by reweighting Bellman targets.
Selective state-adaptive regularization improves offline RL performance.
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the…
FORE evaluates occupancy ratios without requiring Bellman completeness.
New method improves stability of soft FQI for offline RL.
A new estimator combines bootstrapping and rollout methods in RL.
Deep neural nets approximate high-dimensional HJB equations efficiently.
Improved machine learning for reservoir optimization problems.
We address the problem of automatic generation of features for value function approximation. Bellman Error Basis Functions (BEBFs) have been shown to improve the error of policy evaluation with function approximation, with a convergence rate similar to that of value iteration. We propose a simple, fast and robust algor…
The paper introduces Bellman-consistent pessimism to improve offline reinforcement learning without overly pessimistic bias.
Deep Galerkin Method estimates value function for mean-field control problem.
One-step Bellman alignment improves online RL by reducing task mismatch.
A method for calculating multi-portfolio time consistent multivariate risk measures in discrete time is presented. Market models for assets with transaction costs or illiquidity and possible trading constraints are considered on a finite probability space. The set of capital requirements at each time and state is c…
We characterize the value of swing contracts in continuous time as the unique viscosity solution of a Hamilton-Jacobi-Bellman equation with suitable boundary conditions. The case of contracts with penalties is straightforward, and in that case only a terminal condition is needed. Conversely, the case of contracts with …
We study an optimal execution problem in a continuous-time market model that considers market impact. We formulate the problem as a stochastic control problem and investigate properties of the corresponding value function. We find that right-continuity at the time origin is associated with the strength of market impact…
New approach uses PDE learning for faster RL fine-tuning.
Trading strategy mimics optimal control with simple heuristic.
The recently proposed distributional approach to reinforcement learning (DiRL) is centered on learning the distribution of the reward-to-go, often referred to as the value distribution. In this work, we show that the distributional Bellman equation, which drives DiRL methods, is equivalent to a generative adversarial n…
In a discounted reward Markov Decision Process (MDP), the objective is to find the optimal value function, i.e., the value function corresponding to an optimal policy. This problem reduces to solving a functional equation known as the Bellman equation and a fixed point iteration scheme known as the value iteration is u…
In this paper, we extend the jump-diffusion model proposed by Davis and Lleo to include jumps in asset prices as well as valuation factors. The criterion, following earlier work by Bielecki, Pliska, Nagai and others, is risk-sensitive optimization (equivalent to maximizing the expected growth rate subject to a constrai…
Policy evaluation is a key process in Reinforcement Learning (RL). It assesses a given policy by estimating the corresponding value function. When using parameterized value functions, common approaches minimize the sum of squared Bellman temporal-difference errors and receive a point-estimate for the parameters. Kalman…
This paper introduces a set of algorithms for Monte-Carlo Bayesian reinforcement learning. Firstly, Monte-Carlo estimation of upper bounds on the Bayes-optimal value function is employed to construct an optimistic policy. Secondly, gradient-based algorithms for approximate upper and lower bounds are introduced. Finally…
KBB algorithm reduces sample complexity for policy evaluation in general state spaces.
Study optimal consumption and investment strategies with leverage constraints using Epstein-Zin utility.
In this paper we propose and analyze a method based on the Riccati transformation for solving the evolutionary Hamilton-Jacobi-Bellman equation arising from the stochastic dynamic optimal allocation problem. We show how the fully nonlinear Hamilton-Jacobi-Bellman equation can be transformed into a quasi-linear paraboli…
Deep learning method proves convergence for high-dimensional PDEs.
Study shows offline RL under -approximation and partial coverage is harder than previously thought.
Study uses reinforcement learning to optimize portfolios under recursive utility.
Paper addresses underestimation bias in double Q-learning, proposing a method to improve learning performance.
Sequential decision making in the presence of uncertainty and stochastic dynamics gives rise to distributions over state/action trajectories in reinforcement learning (RL) and optimal control problems. This observation has led to a variety of connections between RL and inference in probabilistic graphical models (PGMs)…
In this paper, we consider the problem of online learning of Markov decision processes (MDPs) with very large state spaces. Under the assumptions of realizable function approximation and low Bellman ranks, we develop an online learning algorithm that learns the optimal value function while at the same time achieving ve…
The paper proves well-posedness of nonlocal PDEs related to stochastic control problems.
The paper develops RL methods for optimal switching between multiple states.
Optimizes portfolio in volatile markets with jumps, providing accurate formulas.
Investigates optimal insurance and reinsurance strategies with incomplete market information.
Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot scale to tasks with very high state and action dimensionality such as 3D humanoid l…
We consider a zero-sum stochastic differential controller-and-stopper game in which the state process is a controlled diffusion evolving in a multi-dimensional Euclidean space. In this game, the controller affects both the drift and the volatility terms of the state process. Under appropriate conditions, we show that t…