The paper tackles optimal stopping problems using reinforcement learning and singular control.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A framework for robust exploration in reinforcement learning under ambiguity.
The paper defines and solves time-inconsistent stopping control problems in multi-dimensional diffusion models.
A survey of existing methods for stopping active learning (AL) reveals the needs for methods that are: more widely applicable; more aggressive in saving annotations; and more stable across changing datasets. A new method for stopping AL based on stabilizing predictions is presented that addresses these needs. Furthermo…
Paper optimizes aquaculture feeding and harvesting strategies for profit maximization.
We use martingale and stochastic analysis techniques to study a continuous-time optimal stopping problem, in which the decision maker uses a dynamic convex risk measure to evaluate future rewards. We also find a saddle point for an equivalent zero-sum game of control and stopping, between an agent (the "stopper") who c…
Equivalences are known between problems of singular stochastic control (SSC) with convex performance criteria and related questions of optimal stopping, see for example Karatzas and Shreve [SIAM J. Control Optim. 22 (1984)]. The aim of this paper is to investigate how far connections of this type generalise to a non co…
Study optimal stopping for diffusion processes with unknown primitives, applying RL and martingale methods.
New algorithms control FDX while achieving more power in online multiple testing.
Unified approach to stochastic control, filtering, and stopping using rough paths.
This paper extends stock trading results to include stop-loss orders.
CITE algorithm provides anytime-valid certification of model outputs.
Early stopping method saves up to 75% computation time in policy search tasks.
We consider an optimal stopping problem where a constraint is placed on the distribution of the stopping time. Reformulating the problem in terms of so-called measure-valued martingales allows us to transform the marginal constraint into an initial condition and view the problem as a stochastic control problem; we esta…
Paper introduces a method to control early classification accuracy gaps.
Backdoor attacks on DRL-based traffic controllers cause stop-and-go waves or crashes.
Study optimal stopping in random exploration, deriving HJB and designing a reinforcement learning algorithm.
The paper analyzes optimal retirement timing considering age-dependent mortality risk.
We study the optimal dividend problem for a firm's manager who has partial information on the profitability of the firm. The problem is formulated as one of singular stochastic control with partial information on the drift of the underlying process and with absorption. In the Markovian formulation, we have a 2-dimensio…
In this paper, we present a family of a control-stopping games which arise naturally in equilibrium-based models of market microstructure, as well as in other models with strategic buyers and sellers. A distinctive feature of this family of games is the fact that the agents do not have any exogenously given fundamental…
Optimal reinsurance strategy with fixed cost and exponential preferences.
Extends RL to random stopping times, improving optimization.
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed dynamic optimal control and stopping problems in the existing literature, to stu…
In this paper we study a utility maximization problem with both optimal control and optimal stopping in a finite time horizon. The value function can be characterized by a variational equation that involves a free boundary problem of a fully nonlinear partial differential equation. Using the dual control method, we der…
Study on games with degenerate diffusion matrices, proving value existence and convergence.
Develops a numerical algorithm for stochastic impulse control using regression surrogates.
In this article we study an optimal stopping/optimal control problem which models the decision facing a risk-averse agent over when to sell an asset. The market is incomplete so that the asset exposure cannot be hedged. In addition to the decision over when to sell, the agent has to choose a control strategy which corr…
Improved reinforcement method for optimal control problems.
Optimal retirement timing and consumption under shortfall risk management
We solve the problem of optimal stopping of a Brownian motion subject to the constraint that the stopping time's distribution is a given measure consisting of finitely-many atoms. In particular, we show that this problem can be converted to a finite sequence of state-constrained optimal control problems with additional…
Solves inventory control with unknown demand trend using singular control.
Derives a new formula for optimal stopping problems with exploding derivatives.
Optimal healthcare investment timing in a dynamic model with mortality risk.
Inspired by recent work of P.-L. Lions on conditional optimal control, we introduce a problem of optimal stopping under bounded rationality: the objective is the expected payoff at the time of stopping, conditioned on another event. For instance, an agent may care only about states where she is still alive at the time …
We reveal an interesting convex duality relationship between two problems: (a) minimizing the probability of lifetime ruin when the rate of consumption is stochastic and when the individual can invest in a Black-Scholes financial market; (b) a controller-and-stopper problem, in which the controller controls the drift a…
Using a bondholder who seeks to determine when to sell his bond as our motivating example, we revisit one of Larry Shepp's classical theorems on optimal stopping. We offer a novel proof of Theorem 1 from from \cite{Shepp}. Our approach is that of guessing the optimal control function and proving its optimality with mar…
We study a robust optimal stopping problem with respect to a set $\cP$ of mutually singular probabilities. This can be interpreted as a zero-sum controller-stopper game in which the stopper is trying to maximize its pay-off while an adverse player wants to minimize this payoff by choosing an evaluation criteria from $\…
Consider the problem of a government that wants to reduce the debt-to-GDP (gross domestic product) ratio of a country. The government aims at choosing a debt reduction policy which minimises the total expected cost of having debt, plus the total expected cost of interventions on the debt ratio. We model this problem as…
Study speculative trading using RL with exploratory framework.
We show that deliberately introducing a nested simulation stage can lead to significant variance reductions when comparing two stopping times by Monte Carlo. We derive the optimal number of nested simulations and prove that the algorithm is remarkably robust to misspecifications of this number. The method is applied to…
We study the existence of optimal actions in a zero-sum game between a stopper and a controller choosing a probability measure. This includes the optimal stopping problem for a class of sublinear expectations such as the -expectation. We show that …
New approach solves utility maximization problems using Delta family.
The paper introduces a limit version of multiple stopping options such that the holder selects dynamically a weight function that control the distribution of the payments (benefits) over time. In applications for commodities and energy trading, a control process can represent the quantity that can be purchased by a fix…
Neural networks solve variational inequalities for optimal stopping problems.
We consider the problem of stopping a diffusion process with a payoff functional that renders the problem time-inconsistent. We study stopping decisions of naive agents who reoptimize continuously in time, as well as equilibrium strategies of sophisticated agents who anticipate but lack control over their future selves…
Paper studies early-stopped mirror descent for noisy sparse phase retrieval.
New findings control FDR for online testing methods under positive dependence.
From the Hamilton-Jacobi-Bellman equation for the value function we derive a non-linear partial differential equation for the optimal portfolio strategy (the dynamic control). The equation is general in the sense that it does not depend on the terminal utility and provides additional analytical insight for some optimal…