Improves policies with high certainty, even in small samples.
problem Ensuring new policies are better than the baseline with high probability.
method Leverages powerful safety tests and multiple testing for threshold policies.
result Controls the rate of adopting a worse policy to pre-specified error level.
To overcome the curse of dimensionality and curse of modeling in Dynamic Programming (DP) methods for solving classical Markov Decision Process (MDP) problems, Reinforcement Learning (RL) algorithms are popular. In this paper, we consider an infinite-horizon average reward MDP problem and prove the optimality of the th…
To overcome the curses of dimensionality and modeling of Dynamic Programming (DP) methods to solve Markov Decision Process (MDP) problems, Reinforcement Learning (RL) methods are adopted in practice. Contrary to traditional RL algorithms which do not consider the structural properties of the optimal policy, we propose …
Study optimal treatment assignment policies under strategic agent responses.
problem Learning optimal treatment policies with strategic agents complicates estimation.
method Dynamic model with threshold convergence to mean-field equilibrium, consistent estimator for policy gradient.
result Threshold for treatment assignment converges to mean-field equilibrium threshold under large but finite number of agents.
We discuss the turnpike property for optimal investment and consumption problems. We find there exists a threshold value that determines the turnpike property for investment policy. The threshold value only depends on the Sharpe ratio, the riskless interest rate and the discount rate. We show that if utilities behave a…
Paper shows equivalence between two dividend preference models.
problem Understanding investor and firm preferences for dividends.
method Formulated Epstein-Zin preference, proved equivalence with Maenhout's model.
result Robust dividend policy is equivalent to a threshold strategy based on surplus process.
This article examines arbitrage investment in a mispriced asset when the mispricing follows the Ornstein-Uhlenbeck process and a credit-constrained investor maximizes a generalization of the Kelly criterion. The optimal differentiable and threshold policies are derived. The optimal differentiable policy is linear with …
The trade-off between the cost of acquiring and processing data, and uncertainty due to a lack of data is fundamental in machine learning. A basic instance of this trade-off is the problem of deciding when to make noisy and costly observations of a discrete-time Gaussian random walk, so as to minimise the posterior var…
Study analyzes FIT schemes under market and regulatory uncertainty.
problem Tackles uncertainty in feed-in tariffs and their impact on investment thresholds.
method Uses semi-analytical real options framework to model and compare FIT schemes.
result Increasing regulatory uncertainty lowers investment thresholds for FIT schemes.
Study optimal policies under budget and coverage constraints.
problem Optimal policy learning with budget and coverage constraints.
method Combination of knapsack structure, affine threshold rule, linear programming relaxation, Greedy-Lagrangian (GLC), and rank-and-cut (RC) algorithms.
result GLC closely approximates the optimal solution and achieves near-optimal performance in finite samples; RC is approximately optimal under certain conditions.
Typically, operational risk losses are reported above a threshold. Fitting data reported above a constant threshold is a well known and studied problem. However, in practice, the losses are scaled for business and other factors before the fitting and thus the threshold is varying across the scaled data sample. A report…
Excessively changing policies in many real world scenarios is difficult, unethical, or expensive. After all, doctor guidelines, tax codes, and price lists can only be reprinted so often. We may thus want to only change a policy when it is probable that the change is beneficial. In cases that a policy is a threshold on …
In this research we study a finite horizon optimal purchasing problem for items with a mean reverting price process. Under this model a fixed amount of identical items are bought under a given deadline, with the objective of minimizing the cost of their purchasing price and associated holding cost. We prove that the op…
A learning-based algorithm optimizes admission control in a queuing system.
problem Optimizing admission decisions in a queuing system with unknown parameters.
method Proposes a learning-based dispatching algorithm to minimize regret compared to optimal policies.
result Achieves optimal regret bounds for different scenarios of unknown parameters.
Risk-controlled post-processing optimizes decision policies under risk constraints.
problem Optimizing decision policies with risk constraints for better outcomes.
method Developed a post-processing algorithm that selects a threshold based on fitted fallback policy and score, leveraging tools from algorithmic stability and stochastic processes.
result The post-processed policy achieves precise expected risk control under exchangeability and meets or nearly meets risk budgets while preserving more agreement with the baseline.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
Optimizes state monitoring in Markovian systems with cost constraints.
problem Balancing state queries with prediction costs in Markovian systems.
method Greedy policy and SGD-based learning variant for optimal predict-query tradeoff.
result Greedy policy is suboptimal but performs close to optimal under certain conditions.
We consider a version of the stochastic inventory control problem for a spectrally positive Lévy demand process, in which the inventory can only be replenished at independent exponential times. We show the optimality of a periodic barrier replenishment policy that restocks any shortage below a certain threshold at each…
Develops a support-aware framework for reserve-policy selection in advertising markets.
problem Log-based reserve-price evaluation risks weak support and subgroup harm.
method Support-aware offline decision framework converting logged evidence into certified policies.
result Preserves the best gate-passing policy while eliminating only policies with certified regret.
Percolation on complex networks has been used to study computer viruses, epidemics, and other casual processes. Here, we present conditions for the existence of a network specific, observation dependent, phase transition in the updated posterior of node states resulting from actively monitoring the network. Since tradi…
Even in the face of deteriorating and highly volatile demand, firms often invest in, rather than discard, aging technologies. In order to study this phenomenon, we model the firm's profit stream as a Brownian motion with negative drift. At each point in time, the firm can continue operations, or it can stop and exit th…
Robust policies improve ICU transfer outcomes by predicting patient deterioration.
problem Higher mortality rates for unplanned ICU transfers.
method Markov Decision Process model to predict patient severity and optimize transfer policies.
result Robust policies are more aggressive in transferring patients than nominal policies, improving overall patient care.
Optimal control of reserve assets for stablecoins to maintain peg stability.
problem Balancing immediate liquidity and yield on reserve assets for stablecoin peg maintenance.
method Developed a stochastic model predictive control framework with moment closure for event intensities, incorporating a soft-thresholding structure for rebalancing.
result Optimal policy shifts predictably toward cash as expected outflows intensify or windows lengthen, preserving most bill carry in calm markets and quickly building cash during stress.
Default risk significantly affects the corporate policies of a firm. We develop a model in which a limited liability entity subject to Poisson default shock jointly sets its dividend policy and capital structure to maximize the expected lifetime utility from consumption of risk averse equity investors. We give a comple…
The paper establishes a nearly-sharp statistical threshold for efficient learning in Latent MDPs with separated components.
problem Learning Latent Markov Decision Processes (LMDPs) with separated components.
method The paper considers various notions of separation and establishes a nearly-sharp statistical threshold for efficient learning. It also presents a quasi-polynomial algorithm with time complexity scaling in terms of the statistical threshold under a weaker assumption of separability under the optimal policy, and a near-matching time complexity lower bound under the exponential time hypothesis.
result Establishes a nearly-sharp statistical threshold for efficient learning in Latent MDPs with separated components.
We propose the use of Bayesian networks, which provide both a mean value and an uncertainty estimate as output, to enhance the safety of learned control policies under circumstances in which a test-time input differs significantly from the training set. Our algorithm combines reinforcement learning and end-to-end imita…
Finite-time queue peaks in stochastic networks have logarithmic scaling after geometric thresholds.
problem Queue peak laws in stochastic networks with geometric thresholds.
method Self-normalization mechanism
result Logarithmic scaling of queue peaks after geometric thresholds.
ESRL uses uncertainty quantification to learn safe, optimal policies in offline RL.
problem Challenges in interpreting and measuring uncertainty of learned policies in offline RL.
method Expert-Supervised Reinforcement Learning (ESRL) framework that uses hypothesis testing and posterior distributions.
result The framework can learn safe and optimal policies with theoretical guarantees and independent sample efficiency.
New algorithm tackles unknown utility network resource allocation.
problem Maximizing network utility with unknown agent utilities.
method Modeling as a bandit problem, proposing algorithms for resource allocation.
result Proposed algorithms are optimal when all agents have the same utility.
MTSSL optimizes threshold τ for better semi-supervised learning performance.
problem Optimizing the threshold τ for effective semi-supervised learning.
method Meta-Thresholding approach to optimize τ during training.
result Optimal values of τ are not necessary for achieving similar performance.
The paper tackles reward-relevance in offline RL with sparse decision dynamics.
problem Offline reinforcement learning with sparse decision dynamics and estimation sparsity.
method Reward-filtered least-squares policy evaluation using thresholded lasso.
result The method provides theoretical guarantees with sample complexity dependent on sparse component size.
New bandit model for healthcare intervention planning.
problem Maximizing patient health with limited monitoring resources.
method Developed Collapsing Bandits model and derived optimal policies.
result 3-order-of-magnitude speedup in algorithm performance.
Modern deep learning methods provide effective means to learn good representations. However, is a good representation itself sufficient for sample efficient reinforcement learning? This question has largely been studied only with respect to (worst-case) approximation error, in the more classical approximate dynamic pro…
Gradient-based algorithm improves model performance with triage.
problem Improving model accuracy with human expert involvement.
method Formal characterization of triage, optimal triage policy as deterministic threshold rule, gradient-based algorithm.
result Gradient-based algorithm finds triage policies and models with increasing performance.
The paper studies early stopping methods in linear contextual bandits.
problem Minimizing in-experiment regret and conducting robust post-experiment inferences in contextual bandits.
method The study proposes early stopping rules based on the Opportunity Cost and Threshold Method, using variances of estimators to quantify upper regret bounds.
result The proposed method provides a systematic approach to minimize in-experiment regret and conduct robust post-experiment inferences.
Reward-poisoning attacks can force RL agents to learn bad policies, and we categorize and quantify their feasibility.
problem Reward-poisoning attacks can manipulate RL agents to learn undesirable policies.
method Categorize attacks by infinity-norm constraint, provide thresholds for feasibility, and develop adaptive attack strategies.
result Adaptive reward-poisoning attacks can achieve the nefarious policy in polynomial steps, while non-adaptive attacks require exponential steps.
This paper considers the optimal dividend payment problem in piecewise-deterministic compound Poisson risk models. The objective is to maximize the expected discounted dividend payout up to the time of ruin. We provide a comparative study in this general framework of both restricted and unrestricted payment schemes, wh…
Success conditioning optimizes policies by imitating successful trajectories, solving a trust-region optimization problem.
problem Improving policies through random actions that lead to desired outcomes.
method Success conditioning, which involves collecting and updating policies based on successful trajectories.
result Success conditioning solves a trust-region optimization problem, maximizing policy improvement with a χ2 divergence constraint. Canary optimizes VaR-constrained RL problems with a conservative bound using Cantelli's inequality.
problem Optimizing reinforcement learning policies under VaR constraints in dense cost regimes.
method Employing Cantelli's inequality to create a conservative and smooth bound on VaR constraints based on moments of cost returns. Extending trust-region framework for worst-case bounds on policy improvement and constraint violation.
result Canary reliably satisfies VaR constraints with fewest violations and earliest permanent satisfaction, while maintaining reward competitiveness.
Study on policy testing in MDPs with lower bounds and new algorithm.
problem Deciding if policy value exceeds a threshold with limited samples.
method Derived lower bound, proposed new algorithm, reformulated problem, used policy optimization in reversed MDP.
result New algorithm outperforms existing methods in policy testing.
Efficiently computes optimal policies for Entropic Risk Measures.
problem Optimizing risk-sensitive metrics in MDPs is computationally expensive.
method Uses Entropic Risk Measures and novel structural analysis for efficient computation.
result Achieves strong performance in various decision-making scenarios.
The paper analyzes RLVR's training dynamics, proving convergence depends on aligning update direction with Gradient Gap.
problem Understanding why RLVR works and its limitations.
method Analysis of RLVR's training process at trajectory and token levels, introducing Gradient Gap.
result Convergence depends on aligning update direction with Gradient Gap, with a sharp step-size threshold.
Consider the optimal dividend problem for an insurance company whose uncontrolled surplus precess evolves as a spectrally negative Levy process. We assume that dividends are paid to the shareholders according to admissible strategies whose dividend rate is bounded by a constant. The objective is to find a dividend poli…
New algorithm learns optimal policies with minimal memory and time.
problem Learning optimal policies in discounted MDPs with short burn-in time.
method Variance reduction and adaptive policy switching.
result First regret-optimal model-free algorithm with low burn-in time.
New method estimates optimal dose intervals for personalized treatment.
problem Learning optimal dose intervals from observational data.
method Probability dose interval (PDI) method using DC algorithm.
result Consistent policy with risk converging to best-in-class at root-n rate.
Consider a platform that wants to learn a personalized policy for each user, but the platform faces the risk of a user abandoning the platform if she is dissatisfied with the actions of the platform. For example, a platform is interested in personalizing the number of newsletters it sends, but faces the risk that the u…
This work tackles risk-sensitive deep RL by optimizing policies with variance constraints.
problem Risk and aleatoric uncertainty in deep reinforcement learning.
method Lagrangian and Fenchel dualities to transform the problem into an unconstrained saddle-point policy optimization problem, and an actor-critic algorithm to iteratively update policy, Lagrange multiplier, and Fenchel dual variable.
result The proposed actor-critic algorithm finds a globally optimal policy at a sublinear rate.
Sayer uses implicit feedback to optimize system policies.
problem Leveraging implicit feedback to improve system policies is difficult due to bias and incompleteness.
method Sayer combines randomized exploration and unbiased counterfactual estimators to evaluate and train new policies using implicit feedback.
result Sayer can accurately evaluate and train new policies that outperform existing ones.