Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes optimal consumption with past spending maximum as a reference.
The paper studies scaling limits of hedging prices in financial models.
We study power utility maximization for exponential Lévy models with portfolio constraints, where utility is obtained from consumption and/or terminal wealth. For convex constraints, an explicit solution in terms of the Lévy triplet is constructed under minimal assumptions by solving the Bellman equation. We use a nove…
Study optimal reinsurance and investment strategies under common shocks affecting financial and actuarial markets.
The paper explores solutions to the distributional Bellman equation in reinforcement learning.
New RL approach handles non-exponential discounting for sequential decisions.
In this paper, we consider the problem of optimal investment by an insurer. The insurer invests in a market consisting of a bank account and risky assets. The mean returns and volatilities of the risky assets depend nonlinearly on economic factors that are formulated as the solutions of general stochastic different…
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
Study optimal investment strategies for an insurer in two currency markets.
We study utility maximization for power utility random fields with and without intermediate consumption in a general semimartingale model with closed portfolio constraints. We show that any optimal strategy leads to a solution of the corresponding Bellman equation. The optimal strategies are described pointwise in term…
Sequential decision making in the presence of uncertainty and stochastic dynamics gives rise to distributions over state/action trajectories in reinforcement learning (RL) and optimal control problems. This observation has led to a variety of connections between RL and inference in probabilistic graphical models (PGMs)…
Study solves optimal portfolio selection using HJB equation.
Choquet regularization improves exploration in RL.
Deep neural nets approximate high-dimensional HJB equations efficiently.
This paper concerns the dual risk model, dual to the risk model for insurance applications, where premiums are surplus-dependent. In such a model premiums are regarded as costs, while claims refer to profits. We calculate the mean of the cumulative discounted dividends paid until ruin, if the barrier strategy is applie…
Model quantifies uncertainty's impact on European option prices.
Optimizes portfolios with costs, showing existence of optimal strategies.
New RL algorithm minimizes distributional learning error.
Paper solves investment strategy optimization with deep learning.
The aim of this paper is to construct and analyze solutions to a class of Hamilton-Jacobi-Bellman equations with range bounds on the optimal response variable. Using the Riccati transformation we derive and analyze a fully nonlinear parabolic partial differential equation for the optimal response function. We construct…
We present an approach for pricing European call options in presence of proportional transaction costs, when the stock price follows a general exponential Lévy process. The model is a generalization of the celebrated work of Davis, Panas and Zariphopoulou (1993), where the value of the option is defined as the utility …
Study risk-sensitive market making with entropy regularization for better quote control.
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…
We consider a semilinear parabolic degenerated Hamilton-Jacobi-Bellman (HJB) equation with singularity which is related to a stochastic control problem with fuel constraint. The fuel constraint translates into a singular initial condition for the HJB equation. We first propose a transformation based on a change of vari…
In this paper we propose and analyze a method based on the Riccati transformation for solving the evolutionary Hamilton-Jacobi-Bellman equation arising from the stochastic dynamic optimal allocation problem. We show how the fully nonlinear Hamilton-Jacobi-Bellman equation can be transformed into a quasi-linear paraboli…
The paper solves a complex financial optimization problem using a novel mathematical technique.
Deep Bellman Hedging uses reinforcement learning to optimize financial portfolio hedging.
We study the problem of dynamically trading futures in a regime-switching market. Modeling the underlying asset price as a Markov-modulated diffusion process, we present a utility maximization approach to determine the optimal futures trading strategy. This leads to the analysis of the associated system of Hamilton-Jac…
The paper introduces Bellman-consistent pessimism to improve offline reinforcement learning without overly pessimistic bias.
This paper extends the classical consumption and portfolio rules model in continuous time (Merton 1969, 1971) to the framework of decision-makers with time-inconsistent preferences. The model is solved for different utility functions for both, naive and sophisticated agents, and the results are compared. In order to so…
Optimal contracts are found for agents with quadratic effort costs.
Paper analyzes convergence of Sinkhorn algorithm for discrete probability measures on torus.
The problem of determining the European-style option price in the incomplete market has been examined within the framework of stochastic optimization. An analytic method based on the discrete dynamic programming equation (Bellman equation) has been developed that gives the general formalism for determining the option p…
We solve continuous-time reinforcement learning using distributional Hamilton-Jacobi-Bellman equations.
New method quantifies uncertainty in reinforcement learning models.
New method for off-policy evaluation in POMDPs using future-dependent value functions.
Deep nets solve MDPs without high dimensions.
We study the optimal excess-of-loss reinsurance problem when both the intensity of the claims arrival process and the claim size distribution are influenced by an exogenous stochastic factor. We assume that the insurer's surplus is governed by a marked point process with dual-predictable projection affected by an envir…
Deep learning for HJB PDEs using synthetic data and residual minimization.
We study the structure of a simple dynamic optimization problem consisting of one state and one control variable, from a physicist's point of view. By using an analogy to a physical model, we study this system in the classical and quantum frameworks. Classically, the dynamic optimization problem is equivalent to a clas…
We extend the Deep Galerkin Method (DGM) introduced in Sirignano and Spiliopoulos (2018)} to solve a number of partial differential equations (PDEs) that arise in the context of optimal stochastic control and mean field games. First, we consider PDEs where the function is constrained to be positive and integrate to uni…
New method uses TT approximations to solve HJB equations for efficient sampling.
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the…
Optimizes control of infectious disease spread using stochastic methods.
Bayes optimal algorithm under certain conditions doesn't achieve exponential simple regret.
In this work we investigate the optimal proportional reinsurance-investment strategy of an insurance company which wishes to maximize the expected exponential utility of its terminal wealth in a finite time horizon. Our goal is to extend the classical Cramer-Lundberg model introducing a stochastic factor which affects …
We consider a general class of non-linear Bellman equations. These open up a design space of algorithms that have interesting properties, which has two potential advantages. First, we can perhaps better model natural phenomena. For instance, hyperbolic discounting has been proposed as a mathematical model that matches …