New framework for policy gradient methods in continuous time reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study proves value of non-Markovian games with partial, asymmetric info.
New method uses neural networks for optimal stopping time problems.
Game (Israeli) options in a multi-asset market model with proportional transaction costs are studied in the case when the buyer is allowed to exercise the option and the seller has the right to cancel the option gradually at a mixed (or randomised) stopping time, rather than instantly at an ordinary stopping time. Allo…
American options are studied in a general discrete market in the presence of proportional transaction costs, modelled as bid-ask spreads. Pricing algorithms and constructions of hedging strategies, stopping times and martingale representations are presented for short (seller's) and long (buyer's) positions in an Americ…
American options in a multi-asset market model with proportional transaction costs are studied in the case when the holder of an option is able to exercise it gradually at a so-called mixed (randomised) stopping time. The introduction of gradual exercise leads to tighter bounds on the option price when compared to the …
New method uses randomised signatures for generating financial time series data.
Paper defines saddle points in asymmetric Dynkin games using martingale theory.
We price and hedge American options robustly in continuous time.
We extend previous large deviations results for the randomised Heston model to the case of moderate deviations. The proofs involve the Gärtner-Ellis theorem and sharp large deviations tools.
We propose a randomised version of the Heston model-a widely used stochastic volatility model in mathematical finance-assuming that the starting point of the variance process is a random variable. In such a system, we study the small-and large-time behaviours of the implied volatility, and show that the proposed random…
Randomized exploration in linear bandits achieves optimal regret bounds.
Private machine learning framework using randomised response.
Randomised classifiers outperform deterministic ones in strategic classification.
Unified high-probability regret bounds for online convex optimisation with randomised gradient estimators.
New algorithm improves game learning with randomised optimism.
Improved Bayesian optimisation method using randomised Gaussian process UCB.
We develop a new Monte Carlo variance reduction method to estimate the expectation of two commonly encountered path-dependent functionals: first-passage times and occupation times of sets. The method is based on a recursive approximation of the first-passage time probability and expected occupation time of sets of a Le…
Solves optimal stopping problem with Poisson constraints using jumps.
Discrete time analogues of ergodic stochastic differential equations (SDEs) are one of the most popular and flexible tools for sampling high-dimensional probability measures. Non-asymptotic analysis in the Wasserstein distance of sampling algorithms based on Euler discretisations of SDEs has been recently develop…
Continuous-time optimal stopping solved with deep reinforcement learning
New algorithms improve stopping time for best arm identification.
We consider the optimal double stopping time problem defined for each stopping time by $v(S)=\esssup\{E[ψ(τ_1, τ_2) | \F_S], τ_1, τ_2 \geq S \}$. Following the optimal one stopping time problem, we study the existence of optimal stopping times and give a method to compute them. The key point is the construction of …
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
Two modified tests improve the reliability of evaluating explanation methods.
We consider a zero-sum continuous time stopping game in which the pay-off is revealed in the maximum of the two stopping times instead of the minimum, which is the case in Dynkin games.
Solves optimal stopping for Gauss-Markov bridges using time-space transformation.
We use probabilistic methods to characterise time dependent optimal stopping boundaries in a problem of multiple optimal stopping on a finite time horizon. Motivated by financial applications we consider a payoff of immediate stopping of "put" type and the underlying dynamics follows a geometric Brownian motion. The op…
In this paper, we propose several "measurements" of the "non-stopping timeness" of ends g of previsible sets, such that g avoids stopping times, in an ambiant filtration. We then study several explicit examples, involving last passage times of some remarkable martingales.
In this work we consider optimal stopping problems with conditional convex risk measures called optimised certainty equivalents. Without assuming any kind of time-consistency for the underlying family of risk measures, we derive a novel representation for the solution of the optimal stopping problem. In particular, we …
The paper tackles optimal stopping problems using reinforcement learning and singular control.
Early stopping method saves up to 75% computation time in policy search tasks.
Paper solves a complex stopping problem using regularization and HJB equations.
Method calculates Parisian stopping times and option prices using Markov chains.
Study optimal stopping times under regime-switching models with constraints.
Numerous kinds of uncertainties may affect an economy, e.g. economic, political, and environmental ones. We model the aggregate impact by the uncertainties on an economy and its associated financial market by randomised mixtures of Lévy processes. We assume that market participants observe the randomised mixtures only …
Inspired by Strotz's consistent planning strategy, we formulate the infinite horizon mean-variance stopping problem as a subgame perfect Nash equilibrium in order to determine time consistent strategies with no regret. Equilibria among stopping times or randomized stopping times may not exist. This motivates us to cons…
We design a randomised parallel version of Adaboost based on previous studies on parallel coordinate descent. The algorithm uses the fact that the logarithm of the exponential loss is a function with coordinate-wise Lipschitz continuous gradient, in order to define the step lengths. We provide the proof of convergence …
New method models stopping times that can be equal with non-zero probability.
We show, under weaker assumptions than in the previous literature, that a perpetual optimal stopping game always has a value. We also show that there exists an optimal stopping time for the seller, but not necessarily for the buyer. Moreover, conditions are provided under which the existence of an optimal stopping time…
This paper considers a time-inconsistent stopping problem in which the inconsistency arises from non-constant time preference rates. We show that the smooth pasting principle, the main approach that has been used to construct explicit solutions for conventional time-consistent optimal stopping problems, may fail under …
We consider two-player non-zero-sum stopping games in discrete time. Unlike Dynkin games, in our games the payoff of each player is revealed after both players stop. Moreover, each player can adjust her own stopping strategy according to the other player's action. In the first part of the paper, we consider the game wh…
This paper extends results of Mortimer and Williams (1991) about changes of probability measure up to a random time under the assumptions that all martingales are continuous and that the random time avoids stopping times. We consider locally absolutely continuous measure changes up to a random time, changes of probabil…
Existence of strong randomized equilibria in mean-field games with common noise.
Early stopping improves sample quality in latent diffusion models.
Motivated by the industry practice of pairs trading, we study the optimal timing strategies for trading a mean-reverting price spread. An optimal double stopping problem is formulated to analyze the timing to start and subsequently liquidate the position subject to transaction costs. Modeling the price spread by an Orn…
In the standard models for optimal multiple stopping problems it is assumed that between two exercises there is always a time period of deterministic length , the so called refraction period. This prevents the optimal exercise times from bunching up together on top of the optimal stopping time for the one-exercise c…
A framework for robust exploration in reinforcement learning under ambiguity.