New objective function preserves Bellman's principle for policy gradient.
problem Lack of objective function capturing Bellman's principle optimally.
method Proposed a new objective function and its gradient.
result Preserves Bellman's principle of optimality in policy gradient.
Richard Bellman's Principle of Optimality, formulated in 1957, is the heart of dynamic programming, the mathematical discipline which studies the optimal solution of multi-period decision problems. In this paper, we look at the main trading principles of Jesse Livermore, the legendary stock operator whose method was pu…
Paper introduces dynamic strategies for multi-period investment models.
problem Optimizing investment strategies over multiple periods with risk and return considerations.
method Developed a Bellman principle for discrete time multi-period mean-variance models, leading to dynamic optimal strategies and efficient frontiers.
result Dynamic optimal strategies can achieve higher returns with lower risk compared to the 1/n strategy.
Unified approach combines reward maximization and empowerment for RL.
problem Combining reward maximization and empowerment for reinforcement learning.
method Unified Bellman optimality principle for empowered reward maximization.
result Unified approach leads to improved initial and competitive final performance.
New method solves continuous time mean-variance model for consistent investment strategy.
problem Time-consistent optimal strategy for continuous time mean-variance model.
method Developed a new Bellman principle method.
result Obtained a time-consistent dynamic optimal strategy.
Optimizes control of infectious disease spread using stochastic methods.
problem Optimizing control of highly infectious diseases like COVID-19.
method Reformulated Hamilton-Jacobi-Bellman equation as stochastic minimum principle, leading to forward-backward stochastic differential equations.
result Numerous numerical solutions presented under various scenarios.
Polynomial-time RL algorithm for constant actions under linear Bellman completeness.
problem Efficient online reinforcement learning with few actions.
method Polynomial-time algorithm based on linear function approximation.
result First computationally efficient algorithm for RL with constant actions under linear Bellman completeness.
Market makers optimize trading with a new implicit scheme for complex inequalities.
problem Optimizing trading in a limit order book with stochastic and impulse control.
method Implicit numerical scheme coupled with policy iteration algorithm.
result Convergence to the unique viscosity solution of the HJBQVI.
This paper surveys some recent results on existence, uniqueness and removable singularities for fully nonlinear differential equations on manifolds. The discussion also treats restriction theorems and the strong Bellman principle.
Choosing a portfolio of risky assets over time that maximizes the expected return at the same time as it minimizes portfolio risk is a classical problem in Mathematical Finance and is referred to as the dynamic Markowitz problem (when the risk is measured by variance) or more generally, the dynamic mean-risk problem. I…
One-step Bellman alignment improves online RL by reducing task mismatch.
problem Online RL struggles with task similarity defined by rewards or transitions.
method One-step Bellman alignment and re-weighted targeting (RWT) to correct task mismatch.
result Regret bounds show task shift complexity, not target MDP, affects performance.
Develops deep learning methods for solving S-shaped utility maximisation problems.
problem Optimizing portfolios with S-shaped utility and random benchmarks.
method Uses deep learning and duality methods to solve the Hamilton-Jacobi-Bellman equation and adjoint equation.
result Demonstrates the accuracy of deep learning methods for non-concave utility maximisation problems.
Improved machine learning for reservoir optimization problems.
problem Optimizing control in high-dimensional storage problems.
method Modified dynamic programming algorithm with neural networks for Bellman values and conditional cuts.
result Neural networks outperform classical feedforward networks in estimating Bellman values.
A method for calculating multi-portfolio time consistent multivariate risk measures in discrete time is presented. Market models for d assets with transaction costs or illiquidity and possible trading constraints are considered on a finite probability space. The set of capital requirements at each time and state is c…
A new estimator combines bootstrapping and rollout methods in RL.
problem Combining strengths of bootstrapping and rollout methods in RL.
method Subgraph Bellman operators and fixed point solving.
result Upper bound on error approaches optimal TD variance with additional term.
The paper analyzes convergence of neural SDEs as sample size increases.
problem Understanding the limiting behavior of neural SDEs as sample size grows.
method Analyzes Hamilton-Jacobi-Bellman equation and uses stochastic maximum principle.
result Convergence of minima and optimal parameters of neural SDEs as sample size increases.
Paper introduces v-CMC linking causality and utility.
problem Linking causality and utility for value theory.
method Developed a new causal independence principle (v-CMC) and proved its equivalence.
result Equivalence of local, global, and decomposition versions of v-CMC.
Paper tackles lifetime ruin with hedge funds and high-watermark fees, considering drift uncertainty.
problem Lifetime ruin problem with hedge funds and high-watermark fees under drift uncertainty.
method Employed the stochastic Perron's method to characterize the value function as the unique viscosity solution to the HJB equation.
result Characterized the value function as the unique viscosity solution without resorting to the proof of dynamic programming principle.
We apply stochastic Perron's method to a singular control problem where an individual targets at a given consumption rate, invests in a risky financial market in which trading is subject to proportional transaction costs, and seeks to minimize her probability of lifetime ruin. Without relying on the dynamic programming…
This work addresses time inconsistency in risk measures and develops a dynamic programming principle for risk minimization problems.
problem Time inconsistency in optimized certainty equivalents (OCEs) risk measures.
method Enlargement of state space to achieve a substitute for time consistency, derivation of dynamic programming principle.
result Characterization of the value function via viscosity solutions of Hamilton--Jacobi--Bellman--Issacs equations.
Unified theory of θ-expectations derived from chaotic dynamics.
problem Non-convex stochastic control problems outside G-expectations.
method Spectral theory of transfer operators for uniformly hyperbolic flows, viscosity solutions to HJB equations.
result Affine Hessian, non-convex gradient structure of θ-expectation. In this paper, we study an insurer's reinsurance-investment problem under a mean-variance criterion. We show that excess-loss is the unique equilibrium reinsurance strategy under a spectrally negative Lévy insurance model when the reinsurance premium is computed according to the expected value premium principle. Furthe…
We study an optimal execution problem in a continuous-time market model that considers market impact. We formulate the problem as a stochastic control problem and investigate properties of the corresponding value function. We find that right-continuity at the time origin is associated with the strength of market impact…
We consider the value function originating from an expected utility maximization problem with finite fuel constraint and show its close relation to a nonlinear parabolic degenerated Hamilton-Jacobi-Bellman (HJB) equation with singularity. On one hand, we give a so-called verification argument based on the dynamic progr…
We provide a dynamic programming principle for stochastic optimal control problems with expectation constraints. A weak formulation, using test functions and a probabilistic relaxation of the constraint, avoids restrictions related to a measurable selection but still implies the Hamilton-Jacobi-Bellman equation in the …
New method uses neural networks to solve complex PDEs from optimal control theory.
problem Solving high-dimensional Hamilton-Jacobi-Bellman PDEs.
method Iterative diffusion optimization techniques, focusing on path measures and divergences.
result Favourable properties of log-variance divergence for Monte Carlo estimators.
Study optimal consumption and investment strategies with leverage constraints using Epstein-Zin utility.
problem Optimal portfolio choice under leverage constraints and Epstein-Zin utility.
method Established viscosity solution to HJB equation, demonstrated smoothness, characterized optimal strategies, derived explicit solutions.
result Explicit solutions for optimal consumption and investment strategies under leverage constraints.
A new option pricing model handles non-constant risk aversion and transaction costs.
problem Deriving a pricing model for options with varying risk aversion.
method Developed a transformation method to solve the penalized nonlinear PDE and used finite difference discretization.
result Derived bounds on option prices and proposed a numerical scheme.
Direct and indirect RL methods classified and compared.
problem Classifying RL methods for sequential decision making.
method Direct RL solves optimal policy directly, indirect RL solves Bellman equation.
result Direct and indirect RL methods are equivalent and can be unified.
The free energy functional has recently been proposed as a variational principle for bounded rational decision-making, since it instantiates a natural trade-off between utility gains and information processing costs that can be axiomatically derived. Here we apply the free energy principle to general decision trees tha…
We study the structure of a simple dynamic optimization problem consisting of one state and one control variable, from a physicist's point of view. By using an analogy to a physical model, we study this system in the classical and quantum frameworks. Classically, the dynamic optimization problem is equivalent to a clas…
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
problem The Bellman error is a poor proxy for the accuracy of the value function.
method Study of the Bellman equation as a surrogate objective for value prediction accuracy.
result The magnitude of the Bellman error is only weakly related to the distance to the true value function, even with all state-action pairs.
A framework for goal-based investing with penalties for fund transfers.
problem Investors' mental accounting and multiple investment goals.
method Continuous-time portfolio selection with mental costs and penalties.
result The value function is the unique solution to a complex system of equations.
This paper optimizes DC pension plan investments using O-U process and loan.
problem Optimizing investment strategy for DC pension plans under specific market conditions.
method Dynamic programming and Hamilton-Jacobi-Bellman equation to derive optimal investment strategy.
result Explicit expression for optimal investment strategy derived.
New method stabilizes FQE by reweighting Bellman targets.
problem Stability guarantees for FQE often rely on Bellman completeness, which can fail with function approximation.
method Proposes stationary-weighted FQE, reweighting Bellman targets by stationary target-to-behavior density ratio.
result Proves finite-sample linear convergence to stationary projected Bellman fixed point without Bellman completeness.
We extend the stochastic Perron method to analyze the framework of stochastic target games, in which one player tries to find a strategy such that the state process almost surely reaches a given target no matter which action is chosen by the other player. Within this framework, our method produces a viscosity sub-solut…
The paper analyzes optimal consumption with past spending maximum as a reference.
problem Optimal consumption with past spending maximum as a reference.
method Path-dependent exponential utility, Hamilton-Jacobi-Bellman (HJB) equation, dual transform, smooth-fit principle.
result Closed-form solutions for optimal investment and consumption strategies in each region.
A new method calibrates value predictions in offline RL to improve reliability.
problem Difficulty in long-horizon value prediction in offline reinforcement learning.
method Bellman calibration, a weak reliability criterion, and Iterated Bellman Calibration.
result Finite-sample guarantees show that Bellman calibration error is controlled at nonparametric rates.
Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
problem Exponential gap between upper and lower bounds in risk-sensitive RL.
method Identified and addressed deficiencies in existing algorithms and analysis; developed novel analysis and exploration mechanism.
result Improved regret upper bounds over existing ones.
New Bellman error estimator improves offline model selection performance.
problem Selecting the best policy from logged data using mean squared Bellman error.
method Developed a more accurate estimator of MSBE and analyzed conditions for successful OMS.
result New estimator achieves impressive offline model selection performance on diverse tasks.
The paper proposes a principle for dynamically adjusting the granularity of reinforcement learning abstractions.
problem Lack of general principles for dynamically adjusting the granularity of reinforcement learning abstractions.
method The paper proposes a principle based on rate-distortion theory, formalized through a performance certificate decomposing value error into learning and abstraction error bounds.
result Soft state-action abstractions can achieve near-optimal performance under substantial lossy compression of state and action information.
The paper explores solutions to the distributional Bellman equation in reinforcement learning.
problem Distributional reinforcement learning considers complete return distributions, not just expected returns.
method Study existence and uniqueness of solutions to general distributional Bellman equations, linking them to multivariate affine equations.
result Any solution to a distributional Bellman equation can be derived from a multivariate affine distributional equation.
Study optimality in safety-constrained Markov decision processes using asynchronous value iteration and modified Q-learning.
problem Optimality in safety-constrained Markov decision processes with multichain structure.
method Formulated as a zero-sum game, constructed asynchronous value iteration scheme and modified Q-learning algorithm.
result Resolved Bellman's principle of optimality for multichain Markov decision processes and provided learning algorithms.
Study on LOB dynamics using mean-field game theory.
problem Modeling liquidity dynamics in limit order books.
method Mean-field stochastic differential equation and control problem formulation.
result Equilibrium density function of LOB can be derived.
In this paper, we adapt stochastic Perron's method to analyze a stochastic target problem with unbounded controls in a jump diffusion set-up. With this method, we construct a viscosity sub-solution and super-solution to the associated Hamiltonian-Jacobi-Bellman (HJB) equations. Under comparison principles, uniqueness o…
We obtain the classical Hanner inequalities by the Bellman function method. These inequalities give sharp estimates for the moduli of convexity of Lebesgue spaces. Easy ideas from differential geometry help us to find the Bellman function using neither "magic guesses" nor calculations.
Paper studies offline RL with linear approx, focusing on inherent Bellman error.
problem Offline RL with linear approx, focusing on inherent Bellman error.
method Algorithm that succeeds under single-policy coverage condition, leveraging inherent Bellman error.
result Algorithm yields first known guarantee under single-policy coverage, even for linear Bellman completeness.
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…