New reinforcement learning method for multi-armed bandits with special stop actions.
problem Optimizing actions in episodes with feedback and stop actions.
method Episodic Multi-armed Bandits (eMAB) and FeedBack Adaptive Learning (FeedBAL).
result Regret bound with logarithmic and polynomial dependence on episode length.
Discrete-time games reveal payoffs after both players stop, leading to new equilibrium strategies.
problem Non-zero-sum stopping games with delayed payoff revelation.
method Analyzes simultaneous and sequential stopping strategies, proving Nash equilibria in both cases.
result Existence of Nash equilibria in mixed and pure stopping strategies.
Unified stopping rules ensure accurate policies in contextual learning.
problem Stopping data collection to ensure accurate policies in personalized decision problems.
method Developed unified stopping rules based on GLR statistics for pairwise action comparisons.
result Unified stopping rules achieve target precision with fewer samples than benchmarks.
Study optimal stopping in random exploration, deriving HJB and designing a reinforcement learning algorithm.
problem Optimal stopping problem in continuous time with random exploration.
method Transformed optimal stopping to optimal control problem, derived HJB equation, designed reinforcement learning algorithm.
result Convergence rate of policy iteration and comparison to classical optimal stopping.
The paper solves IRL for Bayesian stopping time problems.
problem Identifying optimal actions in Bayesian stopping time problems.
method Novel IRL framework using Bayesian revealed preferences.
result Identifies optimality and constructs cost function estimates.
Study optimal stopping for diffusion processes with unknown primitives, applying RL and martingale methods.
problem Optimal stopping for diffusion processes with unknown model primitives.
method Continuous-time reinforcement learning framework, variational inequality formulation, stochastic optimal control, entropy regularizer, semi-analytical optimal Bernoulli distribution, policy improvement theorem, policy iterations.
result Demonstrated high accuracy in learning value functions and characterizing free boundaries for various optimal stopping problems.
We study the existence of optimal actions in a zero-sum game infτsupPEP[Xτ] between a stopper and a controller choosing a probability measure. This includes the optimal stopping problem infτE(Xτ) for a class of sublinear expectations E(⋅) such as the G-expectation. We show that …
Equivalences are known between problems of singular stochastic control (SSC) with convex performance criteria and related questions of optimal stopping, see for example Karatzas and Shreve [SIAM J. Control Optim. 22 (1984)]. The aim of this paper is to investigate how far connections of this type generalise to a non co…
Deep Q-Learning models optimal exercise strategies for option-type products.
problem Modeling optimal exercise strategies for option-type products.
method Reinforcement learning approach using deep neural networks to approximate the Q-function.
result Pricing the contract at inception and deriving bounds on the option price.
Proposes data-driven methods for estimating conditional expectations.
problem Estimating conditional expectations when underlying density is unknown.
method Data-driven techniques to directly estimate conditional expectations from training data.
result Extends data-driven method to solve nonlinear equations in stochastic optimization.
A new method learns to stop with minimal data, outperforming traditional approaches.
problem Optimal stopping problems with unknown distributions.
method Explore-then-exploit approach with logarithmic exploration phase.
result Performance comparable to full information DP solution with minimal exploration.
Backdoor attacks on DRL-based traffic controllers cause stop-and-go waves or crashes.
problem Vulnerability of DRL-based traffic controllers to machine learning attacks.
method Developed a trigger design methodology based on traffic physics principles.
result Backdoored models can cause stop-and-go traffic waves or AV crashes when triggered.
Solves optimal stopping problem with Poisson constraints using jumps.
problem Optimal stopping with Poisson constraints and jumps.
method Penalized backward stochastic differential equation (PBSDE) with jumps, decomposition method based on Jacod-Pham, comparison theorem of BSDEs with jumps.
result Solves American option pricing in nonlinear markets with Poisson constraints.
Neural networks optimize stopping boundaries in financial instruments.
problem Optimizing stopping boundaries in financial instruments.
method Deep neural networks and empirical risk minimization for parameterizing stopping boundaries.
result Proved existence of stopping boundary under natural assumptions.
New method stops experiments early for harm in diverse groups.
problem Early stopping of experiments for harmful treatment effects in diverse populations.
method Causal machine learning approach (CLASH) for early stopping.
result CLASH effectively stops experiments early for harmful treatment effects in diverse groups.
Optimizes timing for buying and selling assets with a trailing stop.
problem Timing optimal buy and sell points for assets with a trailing stop.
method General linear diffusion framework, optimal double stopping problem, numerical method.
result Proves the optimality of using a sell limit order in conjunction with the trailing stop.
A survey of existing methods for stopping active learning (AL) reveals the needs for methods that are: more widely applicable; more aggressive in saving annotations; and more stable across changing datasets. A new method for stopping AL based on stabilizing predictions is presented that addresses these needs. Furthermo…
In this work we consider optimal stopping problems with conditional convex risk measures called optimised certainty equivalents. Without assuming any kind of time-consistency for the underlying family of risk measures, we derive a novel representation for the solution of the optimal stopping problem. In particular, we …
Bayesian analysis optimizes stop-loss thresholds based on drawdown distributions.
problem Arbitrary stop-loss levels in financial strategies.
method Bayesian analysis of drawdown distributions.
result Systematic selection of optimal stop-loss thresholds.
DO-IQS recovers optimal stopping region from expert trajectories, addressing specific challenges.
problem Recovering optimal stopping region from expert trajectories with unknown gain functions.
method Dynamics-Aware Offline Inverse Q-Learning incorporating temporal information and confidence-based oversampling.
result Demonstrated performance on real and artificial data, including optimal intervention for critical events.
The paper tackles optimal stopping problems using reinforcement learning and singular control.
problem Continuous-time and state-space optimal stopping problems.
method Formulated as a singular control problem with randomized stopping times and penalized cumulative residual entropy.
result Identified unique optimal exploratory strategy through dynamic programming.
New algorithm solves complex stopping problems with robust optimization.
problem Solving complex stochastic optimal stopping problems.
method Simulation-based robust optimization with exact reformulation as a zero-one bilinear program.
result Developed polynomial-time heuristics and algorithms for practical solution.
Study optimal stopping with random maturity under nonlinear expectations.
problem Optimal stopping problem with random maturity under nonlinear expectation.
method Analysis of optimal stopping problem with random maturity under a nonlinear expectation with respect to a weakly compact set of mutually singular probabilities.
result The optimal stopping problem can be viewed as a discretionary stopping problem for a player who can influence both drift and volatility.
Study examines whether biased training data lead to biased algorithms.
problem Tackles the issue of bias in algorithmic decision-making.
method Examines a setting where biased human decision-makers affect training data and labels.
result Clarifies conditions for bias reversal in algorithmic decision rules.
Improved reinforcement learning with emergency stops.
problem Reducing exploration in reinforcement learning.
method Emergency stop mechanisms to reduce sample complexity.
result Significant improvement in sample complexity and speed.
The strategy of early stopping is a regularization technique based on choosing a stopping time for an iterative algorithm. Focusing on non-parametric regression in a reproducing kernel Hilbert space, we analyze the early stopping strategy for a form of gradient-descent applied to the least-squares loss function. We pro…
Solves optimal stopping for Gauss-Markov bridges using time-space transformation.
problem Optimal stopping problem of a Gauss-Markov bridge.
method Time-space transformation approach, Picard iteration algorithm.
result Lipschitz continuity of the optimal stopping boundary and its characterization.
New algorithms improve stopping time for best arm identification.
problem Efficiently identifying the best alternative in experiments.
method Proposed algorithms with exponential-tailed stopping time.
result Proved that some algorithms never stop, leading to new methods.
Solves optimal stopping problem for financial technical analysis.
problem Optimal stopping problem for technical analysis models.
method Wide-class dynamics modeling support/resistance lines.
result Solution to optimal stopping problem for technical analysis.
The paper solves recursive optimal stopping problems in stock trading.
problem Optimal stopping in recursive optimal stopping problems with applications to stock trading.
method Introduced a class of recursive optimal stopping problems and showed well-posedness in a Markovian setting. Determined optimal stopping rules in stock trading models.
result The value function is the unique solution to a fixed point problem and an optimal stopping time exists.
Continuous-time optimal stopping solved with deep reinforcement learning
problem Optimal stopping problems in continuous time
method CARLOS (Continuous-time Adaptive Reinforcement Learning for Optimal Stopping)
result Higher prices than existing Bermudan solvers, approaching American upper bound
Study proposes a stopping criterion for active learning based on error stability.
problem Improving predictive performance in active learning by adaptively annotating samples.
method Proposes a stopping criterion based on error stability for Bayesian active learning.
result Demonstrates the proposed criterion stops active learning at the appropriate timing for various models and datasets.
Dynamic programming for optimal stopping under distribution constraints.
problem Optimal stopping with distributional constraints.
method Reformulating as measure-valued martingales and stochastic control problem.
result Established dynamic programming principle.
Adaptive rule improves kernel-based gradient descent performance.
problem Improving convergence speed of kernel-based gradient descent algorithms.
method Empirical effective dimension for stopping rule, learning theory analysis, integral operator approach.
result Optimal learning rates and iteration bounds for KGD with adaptive stopping rule.
Dual martingales improve primal optimal stopping problem efficiency.
problem Optimal stopping problem in the primal formulation.
method Investigation of dual martingales to improve primal methods.
result Accurate dual martingale approximations reduce primal problem variance.
Early stopping improves logistic regression's calibration and consistency in high dimensions.
problem Improving the statistical performance of gradient descent in overparameterized logistic regression.
method Investigates the effects of early stopping on gradient descent in logistic regression.
result Early-stopped gradient descent is well-calibrated and statistically consistent, while asymptotic gradient descent is not.
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
problem Optimal stopping problems with finite-time horizon and state-dependent discounting.
method Linear diffusion process, time-homogeneous gain function, fine regularity properties, continuity and strict monotonicity proof.
result Proves continuity and strict monotonicity of the optimal stopping boundary under mild assumptions.
This paper analyzes the problem of starting and stopping a Cox-Ingersoll-Ross (CIR) process with fixed costs. In addition, we also study a related optimal switching problem that involves an infinite sequence of starts and stops. We establish the conditions under which the starting-stopping and switching problems admit …
Early stopping improves sample quality in latent diffusion models.
problem Latent diffusion models degrade sample quality with conventional early stopping.
method Analyzed the interaction between latent dimension and stopping time under Gaussian framework.
result Lower-dimensional representations benefit from earlier termination, higher-dimensional spaces require later stopping.
Study on randomized algorithms for optimal stopping problems.
problem Optimal stopping problems in randomized algorithms.
method Forward and backward Monte Carlo based optimisation algorithms.
result Proved convergence of the proposed algorithms and derived convergence rates.
Combines proportional and stop-loss reinsurance for insurer and reinsurer.
problem Addressing conflicting interests between insurer and reinsurer.
method Introduces proportional-stop-loss reinsurance using balanced loss function.
result Maximizes expected surplus for both insurer and reinsurer.
The paper proposes a principle for dynamically adjusting the granularity of reinforcement learning abstractions.
problem Lack of general principles for dynamically adjusting the granularity of reinforcement learning abstractions.
method The paper proposes a principle based on rate-distortion theory, formalized through a performance certificate decomposing value error into learning and abstraction error bounds.
result Soft state-action abstractions can achieve near-optimal performance under substantial lossy compression of state and action information.
The paper analyzes early stopping for boosting algorithms using localized Gaussian complexity.
problem Understanding the performance of early stopping in kernel boosting algorithms.
method Direct connection between stopped iterate performance and localized Gaussian complexity of function classes.
result Optimal stopping rules derived for various kernel classes, showing correspondence with practice.
Paper solves a complex stopping problem using regularization and HJB equations.
problem Time-inconsistent mean-variance optimal stopping problem
method Vanishing regularization method to derive HJB equations and prove existence of solutions
result Formally recovers variational inequalities for original problem
The paper studies early stopping methods in linear contextual bandits.
problem Minimizing in-experiment regret and conducting robust post-experiment inferences in contextual bandits.
method The study proposes early stopping rules based on the Opportunity Cost and Threshold Method, using variances of estimators to quantify upper regret bounds.
result The proposed method provides a systematic approach to minimize in-experiment regret and conduct robust post-experiment inferences.
MUSE provides unbiased stopping estimates for optimal problems.
problem Estimating the utility of optimal stopping problems.
method Backward recursive construction of the Multilevel Unbiased Stopping Estimator (MUSE).
result MUSE achieves ε-accuracy with O(1/ε^2) computational cost.
Criterion for stopping conjugacy class enumeration in triangle groups.
problem Enumerating all conjugacy classes in cocompact triangle groups.
method Encoding by P. Dehornoy and T. Pinsky; stopping criterion based on geometric length.
result Stopping criterion for the generation of conjugacy classes in cocompact triangle groups.
The study reveals optimal early stopping behaviors in deep learning models.
problem Understanding optimal early stopping in deep learning models.
method Theoretical analysis of linear models and experimental validation.
result Two distinct behaviors of optimal early stopping time depending on model dimension relative to dataset features.