Richard Bellman's Principle of Optimality, formulated in 1957, is the heart of dynamic programming, the mathematical discipline which studies the optimal solution of multi-period decision problems. In this paper, we look at the main trading principles of Jesse Livermore, the legendary stock operator whose method was pu…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
Paper introduces dynamic strategies for multi-period investment models.
New method solves continuous time mean-variance model for consistent investment strategy.
Empowerment is an information-theoretic method that can be used to intrinsically motivate learning agents. It attempts to maximize an agent's control over the environment by encouraging visiting states with a large number of reachable next states. Empowered learning has been shown to lead to complex behaviors, without …
Optimizes control of infectious disease spread using stochastic methods.
Polynomial-time RL algorithm for constant actions under linear Bellman completeness.
Market makers optimize trading with a new implicit scheme for complex inequalities.
This paper surveys some recent results on existence, uniqueness and removable singularities for fully nonlinear differential equations on manifolds. The discussion also treats restriction theorems and the strong Bellman principle.
Choosing a portfolio of risky assets over time that maximizes the expected return at the same time as it minimizes portfolio risk is a classical problem in Mathematical Finance and is referred to as the dynamic Markowitz problem (when the risk is measured by variance) or more generally, the dynamic mean-risk problem. I…
One-step Bellman alignment improves online RL by reducing task mismatch.
Develops deep learning methods for solving S-shaped utility maximisation problems.
Improved machine learning for reservoir optimization problems.
A method for calculating multi-portfolio time consistent multivariate risk measures in discrete time is presented. Market models for assets with transaction costs or illiquidity and possible trading constraints are considered on a finite probability space. The set of capital requirements at each time and state is c…
A new estimator combines bootstrapping and rollout methods in RL.
Paper introduces v-CMC linking causality and utility.
The paper analyzes convergence of neural SDEs as sample size increases.
We apply stochastic Perron's method to a singular control problem where an individual targets at a given consumption rate, invests in a risky financial market in which trading is subject to proportional transaction costs, and seeks to minimize her probability of lifetime ruin. Without relying on the dynamic programming…
Unified theory of -expectations derived from chaotic dynamics.
In this paper, we study an insurer's reinsurance-investment problem under a mean-variance criterion. We show that excess-loss is the unique equilibrium reinsurance strategy under a spectrally negative Lévy insurance model when the reinsurance premium is computed according to the expected value premium principle. Furthe…
We study an optimal execution problem in a continuous-time market model that considers market impact. We formulate the problem as a stochastic control problem and investigate properties of the corresponding value function. We find that right-continuity at the time origin is associated with the strength of market impact…
We consider the value function originating from an expected utility maximization problem with finite fuel constraint and show its close relation to a nonlinear parabolic degenerated Hamilton-Jacobi-Bellman (HJB) equation with singularity. On one hand, we give a so-called verification argument based on the dynamic progr…
We provide a dynamic programming principle for stochastic optimal control problems with expectation constraints. A weak formulation, using test functions and a probabilistic relaxation of the constraint, avoids restrictions related to a measurable selection but still implies the Hamilton-Jacobi-Bellman equation in the …
New method uses neural networks to solve complex PDEs from optimal control theory.
Study optimal consumption and investment strategies with leverage constraints using Epstein-Zin utility.
A new option pricing model handles non-constant risk aversion and transaction costs.
The free energy functional has recently been proposed as a variational principle for bounded rational decision-making, since it instantiates a natural trade-off between utility gains and information processing costs that can be axiomatically derived. Here we apply the free energy principle to general decision trees tha…
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
We study the structure of a simple dynamic optimization problem consisting of one state and one control variable, from a physicist's point of view. By using an analogy to a physical model, we study this system in the classical and quantum frameworks. Classically, the dynamic optimization problem is equivalent to a clas…
A framework for goal-based investing with penalties for fund transfers.
This paper optimizes DC pension plan investments using O-U process and loan.
New method stabilizes FQE by reweighting Bellman targets.
We extend the stochastic Perron method to analyze the framework of stochastic target games, in which one player tries to find a strategy such that the state process almost surely reaches a given target no matter which action is chosen by the other player. Within this framework, our method produces a viscosity sub-solut…
The paper analyzes optimal consumption with past spending maximum as a reference.
A new method calibrates value predictions in offline RL to improve reliability.
Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
New Bellman error estimator improves offline model selection performance.
The paper proposes a principle for dynamically adjusting the granularity of reinforcement learning abstractions.
The paper explores solutions to the distributional Bellman equation in reinforcement learning.
Study optimality in safety-constrained Markov decision processes using asynchronous value iteration and modified Q-learning.
Study on LOB dynamics using mean-field game theory.
In this paper, we adapt stochastic Perron's method to analyze a stochastic target problem with unbounded controls in a jump diffusion set-up. With this method, we construct a viscosity sub-solution and super-solution to the associated Hamiltonian-Jacobi-Bellman (HJB) equations. Under comparison principles, uniqueness o…
This paper aims to make a new contribution to the study of lifetime ruin problem by considering investment in two hedge funds with high-watermark fees and drift uncertainty. Due to multi-dimensional performance fees that are charged whenever each fund profit exceeds its historical maximum, the value function is expecte…
We obtain the classical Hanner inequalities by the Bellman function method. These inequalities give sharp estimates for the moduli of convexity of Lebesgue spaces. Easy ideas from differential geometry help us to find the Bellman function using neither "magic guesses" nor calculations.
Paper studies offline RL with linear approx, focusing on inherent Bellman error.
Reinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal poli…
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…
In this paper, we study optimal liquidation problems in a randomly-terminated horizon. We consider the liquidation of a large single-asset portfolio with the aim of minimizing a combination of volatility risk and transaction costs arising from permanent and temporary market impact. Three different scenarios are analyze…