Optimal contracts are found for agents with quadratic effort costs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method stabilizes FQE by reweighting Bellman targets.
A contraction analysis improves model-based RL's error recovery.
We characterize the value of swing contracts in continuous time as the unique viscosity solution of a Hamilton-Jacobi-Bellman equation with suitable boundary conditions. The case of contracts with penalties is straightforward, and in that case only a terminal condition is needed. Conversely, the case of contracts with …
New method improves stability of soft FQI for offline RL.
The paper analyzes off-policy TD-learning using generalized Bellman operators and provides finite-sample bounds.
FORE evaluates occupancy ratios without requiring Bellman completeness.
Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a contraction operator, with a -contraction modulus (where is the …
Stochastic differential equation approximation for linear TD(0) under Markovian noise
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…
The paper analyzes optimal investment strategies for life insurance contracts using mean-variance optimization.
The paper studies risk-sensitive MDPs with recursive risk measures.
This paper optimizes reinsurance contracts with belief differences between insurer and reinsurer.
We address the problem of automatic generation of features for value function approximation. Bellman Error Basis Functions (BEBFs) have been shown to improve the error of policy evaluation with function approximation, with a convergence rate similar to that of value iteration. We propose a simple, fast and robust algor…
Optimal linear contracts are possible even with memory in Gaussian settings.
This work is motivated by numerical solutions to Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) associated with combined stochastic and impulse control problems. In particular, we consider (i) direct control, (ii) penalized, and (iii) semi-Lagrangian discretization schemes applied to the HJBQVI proble…
DSPI connects natural policy gradient to policy iteration, proving global convergence.
In the paper portfolio optimization over long run risk sensitive criterion is considered. It is assumed that economic factors which stimulate asset prices are ergodic but non necessarily uniformly ergodic. Solution to suitable Bellman equation using local span contraction with weighted norms is shown. The form of optim…
In this paper, we consider the stochastic iterative counterpart of the value iteration scheme wherein only noisy and possibly biased approximations of the Bellman operator are available. We call this counterpart as the approximate value iteration (AVI) scheme. Neural networks are often used as function approximators, i…
The aim of this paper is to introduce an insurance model allowing reinsurance and dividend payment. Our model deals with several homogeneous contracts and takes into account the legislation regarding the provisions to be justified by the insurance companies. This translates into some restriction on the (maximal) number…
New Q-learning method achieves optimal sample complexity for average-reward problems.
Paper establishes DRL for high-dimensional rewards.
Study optimal futures trading strategies for assets with multiscale central tendency price model.
Value function learning plays a central role in many state-of-the-art reinforcement-learning algorithms. Many popular algorithms like Q-learning do not optimize any objective function, but are fixed-point iterations of some variant of Bellman operator that is not necessarily a contraction. As a result, they may easily …
We study the problem of dynamically trading multiple futures contracts with different underlying assets. To capture the joint dynamics of stochastic bases for all traded futures, we propose a new model involving a multi-dimensional scaled Brownian bridge that is stopped before price convergence. This leads to the analy…
We analyze reinforcement learning algorithms using a distributional approach.
We study a stochastic control approach to managed futures portfolios. Building on the Schwartz 97 stochastic convenience yield model for commodity prices, we formulate a utility maximization problem for dynamically trading a single-maturity futures or multiple futures contracts over a finite horizon. By analyzing the a…
This paper optimizes perpetual contract liquidity by accounting for funding rates.
In this paper, we take up the analysis of a principal/agent model with moral hazard introduced in [17], with optimal contracting between competitive investors and an impatient bank monitoring a pool of long-term loans subject to Markovian contagion. We provide here a comprehensive mathematical formulation of the model …
Method learns statistics of return distributions via neural networks and maximum mean discrepancy.
We study the problem of dynamically trading a futures contract and its underlying asset under a stochastic basis model. The basis evolution is modeled by a stopped scaled Brownian bridge to account for non-convergence of the basis at maturity. The optimal trading strategies are determined from a utility maximization pr…
Optimal dividend strategy with irreversible reinsurance constraints.
We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifically focus on incorporating robustness into a state-of-the-art continuous control RL algorithm call…
We consider the issue of a market maker acting at the same time in the lit and dark pools of an exchange. The exchange wishes to establish a suitable make-take fees policy to attract transactions on its venues. We first solve the stochastic control problem of the market maker without the intervention of the exchange. T…
New method reduces sample complexity for robust reinforcement learning.
Survey of mathematical foundations for reinforcement learning.
In this paper, we consider a problem of contract theory in which several Principals hire a common Agent and we study the model in the continuous time setting. We show that optimal contracts should satisfy some equilibrium conditions and we reduce the optimisation problem of the Principals to a system of coupled Hamilto…
New methods improve stability of Sinkhorn algorithm in machine learning.
New algorithm for risk-sensitive reinforcement learning with natural policy gradients.
In this paper long-run risk sensitive optimisation problem is studied with dyadic impulse control applied to continuous-time Feller-Markov process. In contrast to the existing literature, focus is put on unbounded and non-uniformly ergodic case by adapting the weight norm approach. In particular, it is shown how to com…
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
Linear Q-learning converges to a bounded set without divergence.
A new method calibrates value predictions in offline RL to improve reliability.
Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
The paper analyzes reinsurance strategies in a competitive multi-agent system.
A variable annuity contract with Guaranteed Minimum Withdrawal Benefit (GMWB) promises to return the entire initial investment through cash withdrawals during the contract plus the remaining account balance at maturity, regardless of the portfolio performance. Under the optimal(dynamic) withdrawal strategy of a policyh…
New Bellman error estimator improves offline model selection performance.
The paper explores solutions to the distributional Bellman equation in reinforcement learning.