The paper explores solutions to the distributional Bellman equation in reinforcement learning.
problem Distributional reinforcement learning considers complete return distributions, not just expected returns.
method Study existence and uniqueness of solutions to general distributional Bellman equations, linking them to multivariate affine equations.
result Any solution to a distributional Bellman equation can be derived from a multivariate affine distributional equation.
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
problem The Bellman error is a poor proxy for the accuracy of the value function.
method Study of the Bellman equation as a surrogate objective for value prediction accuracy.
result The magnitude of the Bellman error is only weakly related to the distance to the true value function, even with all state-action pairs.
Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
problem Exponential gap between upper and lower bounds in risk-sensitive RL.
method Identified and addressed deficiencies in existing algorithms and analysis; developed novel analysis and exploration mechanism.
result Improved regret upper bounds over existing ones.
We study utility maximization for power utility random fields with and without intermediate consumption in a general semimartingale model with closed portfolio constraints. We show that any optimal strategy leads to a solution of the corresponding Bellman equation. The optimal strategies are described pointwise in term…
Study solves optimal portfolio selection using HJB equation.
problem Optimal portfolio selection problem.
method Maximal monotone operator method, Banach fixed-point theorem, Fourier transform, monotone operators technique.
result Existence and uniqueness of solution to HJB equation.
Deep neural nets approximate high-dimensional HJB equations efficiently.
problem Approximating solutions to high-dimensional HJB equations.
method Deep neural networks for approximating solutions.
result Deep neural networks can approximate solutions without the curse of dimensionality.
Model quantifies uncertainty's impact on European option prices.
problem Uncertainty in market volatility risk affects option pricing.
method Hamilton-Jacobi-Bellman framework and finite element method.
result Dependence of Delta on uncertainty is nonlinear and varied.
Optimizes portfolios with costs, showing existence of optimal strategies.
problem Risk-sensitive portfolio optimization with transaction costs.
method Log-return i.i.d. framework, Bellman equation analysis.
result Existence of optimal strategies for risk-averse and risk-seeking cases.
New RL algorithm minimizes distributional learning error.
problem Improving distributional reinforcement learning for better error minimization.
method Proposes a new model-based algorithm with theoretical minimax optimality.
result Proves minimax optimality for approximating return distributions.
Paper solves investment strategy optimization with deep learning.
problem Maximizing investor utility with optimal asset allocation.
method Solves PDEs with Deep Galerkin method.
result Deep learning algorithm outperforms finite difference method.
The aim of this paper is to construct and analyze solutions to a class of Hamilton-Jacobi-Bellman equations with range bounds on the optimal response variable. Using the Riccati transformation we derive and analyze a fully nonlinear parabolic partial differential equation for the optimal response function. We construct…
We consider a semilinear parabolic degenerated Hamilton-Jacobi-Bellman (HJB) equation with singularity which is related to a stochastic control problem with fuel constraint. The fuel constraint translates into a singular initial condition for the HJB equation. We first propose a transformation based on a change of vari…
In this paper we propose and analyze a method based on the Riccati transformation for solving the evolutionary Hamilton-Jacobi-Bellman equation arising from the stochastic dynamic optimal allocation problem. We show how the fully nonlinear Hamilton-Jacobi-Bellman equation can be transformed into a quasi-linear paraboli…
The paper solves a complex financial optimization problem using a novel mathematical technique.
problem Optimizing portfolio selection in financial markets.
method Maximal monotone operator method and Riccati transformation.
result Existence and uniqueness of a solution to the transformed parabolic equation in a Sobolev space.
Deep Bellman Hedging uses reinforcement learning to optimize financial portfolio hedging.
problem Optimizing financial portfolio hedging with derivatives and trading frictions.
method Actor-critic reinforcement learning algorithm with continuous state and action spaces.
result Trained model provides optimal hedge for any initial portfolio and market state.
The paper introduces Bellman-consistent pessimism to improve offline reinforcement learning without overly pessimistic bias.
problem Offline reinforcement learning's challenge of discovering good policies without exhaustive exploration.
method Introduces Bellman-consistent pessimism for function approximation, improving sample complexity and adaptability.
result Improves sample complexity by O(d) in the action space finite case, and automatically adapts to bias-variance tradeoff. Optimal contracts are found for agents with quadratic effort costs.
problem Finding optimal contracts in principal-agent problems with quadratic effort costs.
method Modeling the problem using Hamilton-Jacobi-Bellman (HJB) equations and proving the existence of classical solutions.
result Existence of optimal contracts for agents with quadratic effort costs is proven.
The problem of determining the European-style option price in the incomplete market has been examined within the framework of stochastic optimization. An analytic method based on the discrete dynamic programming equation (Bellman equation) has been developed that gives the general formalism for determining the option p…
We solve continuous-time reinforcement learning using distributional Hamilton-Jacobi-Bellman equations.
problem Predicting the distribution of returns in continuous-time, stochastic environments.
method We derive a distributional Hamilton-Jacobi-Bellman equation for Itô diffusions and Feller-Dynkin processes, and propose an algorithm based on a JKO scheme.
result We propose an online control algorithm that can be used to approximately solve the distributional HJB equation.
New method quantifies uncertainty in reinforcement learning models.
problem Quantifying uncertainty over expected cumulative rewards in reinforcement learning.
method Proposes a new uncertainty Bellman equation to more accurately estimate value function variance.
result Our method converges to the true posterior variance over values and improves sample-efficiency.
New method for off-policy evaluation in POMDPs using future-dependent value functions.
problem Curse of horizon in off-policy evaluation for POMDPs.
method Develops future-dependent value functions and minimax learning method.
result PAC result and Bellman completeness for the proposed OPE estimator.
Deep nets solve MDPs without high dimensions.
problem Solving Bellman equations for MDPs in high dimensions.
method Deep neural networks with ReLU activation approximating payoff and transition functions.
result Deep nets can approximate Q-functions in polynomially bounded parameters. Deep learning for HJB PDEs using synthetic data and residual minimization.
problem Solving Hamilton-Jacobi-Bellman PDEs for optimal control problems.
method Gradient-augmented synthetic dataset for supervised learning, residual minimization.
result Improves accuracy and efficiency of deep learning for HJB PDEs.
We study the structure of a simple dynamic optimization problem consisting of one state and one control variable, from a physicist's point of view. By using an analogy to a physical model, we study this system in the classical and quantum frameworks. Classically, the dynamic optimization problem is equivalent to a clas…
New method uses TT approximations to solve HJB equations for efficient sampling.
problem Efficiently sampling from complex probability densities.
method Direct time integration of HJB equations using Tensor Train compression.
result Sample-free, dimensionality-avoiding integration method.
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the…
Optimizes control of infectious disease spread using stochastic methods.
problem Optimizing control of highly infectious diseases like COVID-19.
method Reformulated Hamilton-Jacobi-Bellman equation as stochastic minimum principle, leading to forward-backward stochastic differential equations.
result Numerous numerical solutions presented under various scenarios.
We consider a general class of non-linear Bellman equations. These open up a design space of algorithms that have interesting properties, which has two potential advantages. First, we can perhaps better model natural phenomena. For instance, hyperbolic discounting has been proposed as a mathematical model that matches …
Stochastic differential equation approximation for linear TD(0) under Markovian noise
problem Temporal-difference learning with linear function approximation
method Stochastic differential equation approximation
result Explains the constant-stepsize error floor
This paper surveys some recent results on existence, uniqueness and removable singularities for fully nonlinear differential equations on manifolds. The discussion also treats restriction theorems and the strong Bellman principle.
New RL formulation for maximizing maximum reward in molecule generation.
problem Traditional RL frameworks do not fit real-world applications like drug discovery.
method Formulated a new objective function to maximize maximum reward, derived Bellman equation, introduced operators, and proved convergence.
result Achieved state-of-the-art results in molecule generation.
In this paper, we extend the jump-diffusion model proposed by Davis and Lleo to include jumps in asset prices as well as valuation factors. The criterion, following earlier work by Bielecki, Pliska, Nagai and others, is risk-sensitive optimization (equivalent to maximizing the expected growth rate subject to a constrai…
A new option pricing model handles non-constant risk aversion and transaction costs.
problem Deriving a pricing model for options with varying risk aversion.
method Developed a transformation method to solve the penalized nonlinear PDE and used finite difference discretization.
result Derived bounds on option prices and proposed a numerical scheme.
New approach transfers rewards learned in one environment to reinforcement learning in a new environment.
problem Transfer of rewards learned using inverse reinforcement learning from one environment to a new, different environment.
method Formulate the problem as a joint system of Bellman equations, develop minimax estimators for the target soft-q-function, solve the source and target system of equations jointly. result The coupled approach removes the first-order influence of source Bellman residual error compared to the sequential approach.
A neural network approach solves optimal decumulation problems for pension plans.
problem Optimal asset allocation and withdrawal strategies for DC pension holders.
method Data-driven neural network optimization with customized activation functions.
result The neural network approach learns near-optimal solutions comparable to HJB PDE methods.
Paper introduces stochastic HJB on Jacobi structures.
problem Stochastic analysis on Jacobi manifolds.
method Global stochastic analysis techniques, extending Bismut and Lázaro-Camí work.
result Proposes a stochastic HJB framework.
A model optimizes carbon emission reduction and allowance purchasing for companies.
problem Optimizing carbon emissions and allowance purchasing for companies.
method Established an optimal control model involving two stochastic processes with two control variables, converted into an HJB equation, proved existence and uniqueness of solution.
result Proved the existence and uniqueness of the solution to the HJB equation.
The Noether theorem is extended to stochastic control problems using contact symmetries.
problem Stochastic optimal control problems.
method Exploiting jet bundles and contact geometry, the authors prove the existence of conserved quantities.
result Optimal control problems admit infinitely many conserved quantities in the form of local martingales.
Corrects gaps in a method for optimizing high-frequency trading strategies.
problem Optimizing bid and ask limit order strategies in high-frequency trading.
method Uses an approximation method based on Avellaneda and Stoikov's 2008 article, correcting gaps found in it.
result The main answer in Avellaneda and Stoikov's article remains unchanged despite corrections.
We consider a utility maximization problem for an investment-consumption portfolio when the current utility depends also on the wealth process. Such kind of problems arise, e.g., in portfolio optimization with random horizon or with random trading times. To overcome the difficulties of the problem we use the dual appro…
The paper proves well-posedness of nonlocal PDEs related to stochastic control problems.
problem Characterizing equilibrium strategies and value functions for time-inconsistent stochastic control problems.
method Method of continuity and Banach's fixed point arguments, with Schauder prior estimates.
result Global well-posedness of nonlocal fully nonlinear PDEs with sharp a-priori estimates.
In this paper we investigate a dynamic stochastic portfolio optimization problem involving both the expected terminal utility and intertemporal utility maximization. We solve the problem by means of a solution to a fully nonlinear evolutionary Hamilton-Jacobi-Bellman (HJB) equation. We propose the so-called Riccati met…
Deep-MacroFin uses neural networks to solve complex economic models efficiently.
problem Solving high-dimensional partial differential equations in continuous time economics.
method Leverages deep learning, specifically Multi-Layer Perceptrons and Kolmogorov-Arnold Networks, optimized with HJB equations.
result Offers a more efficient solution (5imes less memory, 40imes fewer FLOPs) for 50D economic models. The paper tackles non-cumulative objectives in reinforcement learning and proposes modifications to existing algorithms.
problem Optimizing objectives that are not naturally expressed as summations of rewards in various fields.
method The paper modifies the Bellman optimality equation to handle non-cumulative objectives by replacing summation with a generalized operation.
result The modified Bellman updates can converge to the globally optimal solution under certain conditions.
The main purpose of this paper is to analyze solutions to a fully nonlinear parabolic equation arising from the problem of optimal portfolio construction. We show how the problem of optimal stock to bond proportion in the management of pension fund portfolio can be formulated in terms of the solution to the Hamilton-Ja…
Develops deep learning methods for solving S-shaped utility maximisation problems.
problem Optimizing portfolios with S-shaped utility and random benchmarks.
method Uses deep learning and duality methods to solve the Hamilton-Jacobi-Bellman equation and adjoint equation.
result Demonstrates the accuracy of deep learning methods for non-concave utility maximisation problems.
A new macroscopic market making model connects market making and optimal execution.
problem Connecting market making and optimal execution problems.
method Using continuous processes for orders, the model bridges the gap between market making and optimal execution.
result Demonstrates the model's effectiveness through various noise and intensity function scenarios.
The paper calibrates SPX and VIX options using optimal transport.
problem Joint calibration of SPX and VIX options or futures.
method Semimartingale optimal transport problem with PDE formulation and dual formulation.
result The model accurately calibrates SPX, VIX options, and futures simultaneously.