Improved online learning for MDPs with changing costs.
problem Online learning in linearly solvable MDPs with changing state costs.
method Following the leader algorithm with logarithmic regret bound.
result Achieved regret of order log^2 T, significantly better than previous bounds.
New algorithm controls linear systems with bandit feedback, achieving optimal regret.
problem Controlling linear systems with bandit feedback under adversarial costs.
method Developed a new algorithm for linear control with memory optimization technique.
result Achieved optimal regret growth proportional to square root of time horizon.
Paper tackles online optimization with memory and competitive control.
problem Minimizing hitting and switching costs in online optimization problems.
method Optimistic Regularized Online Balanced Descent algorithm.
result Achieves a constant, dimension-free competitive ratio.
OBD algorithm optimizes online convex optimization with strong convexity and switching costs.
problem Online convex optimization with strong convexity and switching costs.
method Online Balanced Descent (OBD) algorithm for m m m -strongly convex costs with near-optimal dynamic regret and per-round accuracy for ε ε ε -smooth sequences. result OBD achieves a competitive ratio of 3 + O ( 1 / m ) 3 + O(1/m) 3 + O ( 1/ m ) for m m m -strongly convex costs. Smoothness of value function in affine control problems proven.
problem Regularity of value function in affine optimal control problems.
method Proved continuity and smoothness on open dense subsets without singular minimizers.
result Value function is smooth on an open dense subset of the interior of the attainable set.
Study bank salvage model with stochastic impulse controls to minimize costs.
problem Minimize total cost of saving a bank from default with unpredictable default time.
method Impulse stochastic controls to address the bank's default risk.
result Unique viscosity solution exists for the QVI, with Lipschitz and Holder continuity properties.
New tontine model with transaction costs for retirees.
problem Maximizing consumption and bequest utilities for retirees.
method Formulated as a stochastic and impulse control problem, characterized by viscosity solutions.
result V-shaped transaction region with two stages: smoothing and gambling.
Study on inventory management under uncertainty using smooth ambiguity preference.
problem Managing inventory under Knightian uncertainty with smooth ambiguity preference.
method Demonstrates continuous-time smooth ambiguity as the infinitesimal limit of Kalman-Bucy filtering with recursive robust utility. Solves forward-backward stochastic differential equations with quadratic growth to determine cost function. Derives value function and optimal control policy using variational inequalities and viscosity solutions. Transforms problem into two-dimensional singular control.
result Ambiguity drives decision-makers to act earlier, reducing the continuation region.
We consider the stochastic control problem of a financial trader that needs to unwind a large asset portfolio within a short period of time. The trader can simultaneously submit active orders to a primary market and passive orders to a dark pool. Our framework is flexible enough to allow for price-dependent impact func…
Counterexample shows state-constrained optimal control problems can have Young measure gaps.
problem Existence of Young measure gaps in state-constrained optimal control problems.
method Provided a counterexample for smooth controllable systems state-constrained to the unit ball.
result Gap occurs in a regular setting with non-convex Lagrangian density.
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
problem Finding optimal policies in nonlinear control systems with partial information.
method Policy gradient algorithm designed for nearly linear-quadratic regulators with small Lipschitz nonlinear components.
result Policy gradient algorithm converges to globally optimal policy with linear rate.
The paper optimizes policies constrained to Schur stabilizing controllers using a Newton-type algorithm.
problem Optimizing policies under linear constraints in control systems.
method Newton-type algorithm on a manifold of Schur stabilizing controllers with a Riemannian metric.
result Local convergence guarantees for the Newton-type algorithm without relying on exponential mapping or retractions.
Non-bilinear observations make optimal control harder, showing non-convex costs and non-affine optimal controllers.
problem Optimal control from bilinear observations in linear systems is challenging.
method Analytical and numerical methods to study the non-convex cost-to-go and non-affine optimal controllers.
result The Separation Principle does not hold for bilinear observations, leading to non-convex costs and non-affine optimal controllers.
FavMac maximizes value while controlling cost in multi-label prediction.
problem Value-maximizing predictions with strict cost control in multi-label scenarios.
method FavMac pipeline combining any multi-label classifier with online update mechanism.
result FavMac achieves higher value with strict cost control compared to baselines.
Meta-learning control algorithm with finite-time guarantees for unknown systems.
problem Online control of unknown linear systems with constraints.
method Provable regret guarantees for an iterative control algorithm.
result Regret bounds of O ( T 3 / 4 ) O(T^{3/4}) O ( T 3/4 ) for controller cost and constraint violation. New method for handling multi-dimensional singular controls with jump costs in mean-field problems.
problem Handling jump costs in multi-dimensional singular controls.
method Introducing two-layer parametrisations to interpolate jumps on both distributional and pathwise levels.
result Derivation of a DPP and characterisation of the value function as a minimal super-solution to a quasi-variational inequality.
Efficient algorithm for unknown linear systems with convex costs.
problem Controlling an unknown linear system with stochastic convex costs.
method Optimism in the Face of Uncertainty paradigm.
result Achieves optimal T \sqrt{T} T regret-rate. Study optimizes investment decisions with fixed costs using stochastic control methods.
problem Optimizing irreversible investment decisions with fixed adjustment costs.
method Stochastic impulse control approach, viscosity solutions, quasi-variational inequality.
result Characterization of optimal control and sensitivity analysis in linear case.
Study cost-driven state representation learning for control from partial observations.
problem Learning state representation for control from partial and high-dimensional observations.
method Cost-driven state representation learning via predicting cumulative costs.
result Established finite-sample guarantees for near-optimal representation and controller.
Study learns state representations from observations for control, proving guarantees.
problem Learning state representations from high-dimensional observations for control.
method Cost-driven approach, learning latent state model to predict costs.
result Proves finite-sample guarantees for near-optimal state representation and controller.
Regularization improves robustness of smoothed classifiers.
problem Certifying robustness of smoothed classifiers.
method Regularizing prediction consistency over Gaussian noise.
result Significantly improved certified robustness with less training costs.
Optimal control in changing systems without strong convexity assumptions.
problem Adversarial changes in convex costs for unknown linear systems.
method Non-convex lower confidence bounds and computationally-efficient regret minimization.
result Achieves T \smash{\sqrt{T}} T -regret rate, optimal compared to best stabilizing controller. Optimizes stock execution costs using stochastic control theory.
problem Minimizing execution costs in a market with discrete stock price movements.
method Discrete-time Stochastic Control Theory applied to a market model.
result Developed optimal allocation strategy for executing K stocks within T units.
Schrödinger bridge solved with Weyl calculus for quadratic state cost.
problem Optimal control policy to steer joint state statistics.
method Weyl calculus in quantum mechanics for reaction-diffusion PDEs.
result Explicit Markov kernel for quadratic state cost found.
Study optimal control of diffusion processes with infimum or supremum costs.
problem Optimizing control of a diffusion process with costs dependent on its infimum or supremum.
method Introduced novel integral operators to solve two-dimensional singular control problems.
result Explicit solutions for optimal dividend problem with time-dependent preferences.
Study optimal pairs trading with transaction costs using stochastic control.
problem Finding optimal trade times and shares in pairs trading with proportional costs.
method Singular stochastic control approach to solve a nonlinear quasi-variational inequality.
result Developed a discrete time dynamic programming algorithm to compute transaction regions.
Paper tackles online control of linear systems with unbounded noise.
problem Online control of linear systems under unbounded noise with unknown convex cost functions.
method Developed an algorithm achieving i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) high-probability regret under unbounded noise, and established O ( m p o l y ( log T ) ) O({
m poly} (\log T)) O ( m p o l y ( log T )) regret bound for strongly convex costs and sub-Gaussian noise. result Achieved i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) high-probability regret under unbounded noise, and O ( m p o l y ( log T ) ) O({
m poly} (\log T)) O ( m p o l y ( log T )) regret bound for specific noise and cost conditions. Derandomizing PAC-Bayes bounds for smooth loss functions
problem Derandomizing PAC-Bayes bounds for smooth loss functions
method Exploiting smoothness properties of both the loss and the predictor class
result Bounds for deterministic predictors that involve flatness quantities
We discuss smooth nonlinear control systems with symmetry. For a free and proper action of the symmetry group, the reduction of symmetry gives rise to a reduced smooth nonlinear control system. If the action of the symmetry group is only proper, the reduced nonlinear control system need not be smooth. Using the smooth …
Derives bounds for deterministic predictors using smooth loss functions.
problem Generalizing probabilistic predictors to deterministic ones.
method Exploits smoothness properties of loss and predictor classes, controlling the Jensen gap class through Rademacher complexity.
result Derives bounds for deterministic predictors involving flatness quantities from Jacobians and Hessians.
Central bank optimizes exchange rate interventions to minimize costs.
problem Minimizing costs of exchange rate interventions by a central bank.
method Singular stochastic control problem with bounded variation controls.
result Explicit expression of the value function and optimal control.
New method handles robust and adaptive control of linear systems with non-convex costs.
problem Robust and adaptive control of linear systems with unknown parameters.
method Combining non-asymptotic linear regression, interval prediction, and tree-based planning.
result First end-to-end suboptimality analysis for robust and adaptive MPC with non-convex costs.
In this paper we study mean-field type control problems with risk-sensitive performance functionals. We establish a stochastic maximum principle (SMP) for optimal control of stochastic differential equations (SDEs) of mean-field type, in which the drift and the diffusion coefficients as well as the performance function…
Safe Bayesian optimization reduces HVAC costs by 32%.
problem Optimizing room temperature PID control for energy savings and comfort.
method Safe Contextual Bayesian Optimization.
result 32% reduction in room temperature PID control costs.
The paper bridges stochastic control and deep hedging for European call options with transaction costs.
problem Hedging and pricing European call options with proportional transaction costs.
method Complementary perspectives: stochastic control and deep hedging. Two architectures proposed: NTBN-Delta and WW-NTBN.
result WW-NTBN converges faster, matches no-transaction bands more closely, and generalizes well across transaction cost regimes.
The study sets limits on how well systems can be controlled adaptively.
problem Learning to control unknown linear Gaussian systems with quadratic costs.
method Combining ideas from experiment design, estimation theory, and perturbation bounds of information matrices.
result Regret lower bounds of the order of T \sqrt{T} T in the time horizon T T T accurately capture control-theoretic parameters. Energy-based model learns cost functions from expert demonstrations for optimal control.
problem Learning unknown cost functions from expert demonstrations for optimal control.
method Maximum likelihood estimation via analysis by synthesis, combining Langevin dynamics with optimization and cooperative learning.
result The method can learn suitable cost functions for optimal control tasks.
Paper solves a control problem with robust methods.
problem Monotone mean-variance problems with stochastic coefficients.
method Finding saddle point through BSDEs with unbounded coefficients.
result Optimal control and value match mean-variance problems.
New control theory shows neural networks can be sparsely active over time.
problem Optimizing neural networks for long-time control with sparsity constraints.
method Proving optimal controls vanish after a positive time and providing a stability estimate.
result Optimal controls for ℓ 1 \ell^1 ℓ 1 -penalized neural ODEs are sparsely active over time. New algorithm achieves logarithmic regret for adversarial online control.
problem Online linear-quadratic control in systems with adversarial disturbances.
method Characterization of optimal offline control law, reduced to online learning with approximate advantage functions.
result First algorithm with logarithmic regret for arbitrary adversarial disturbance sequences.
New algorithm reduces decision switching in dynamic environments.
problem Online learning with memory and non-stationary environments.
method Dynamic policy regret, novel ensemble approach, meta-base decomposition.
result Proves optimal dynamic policy regret for memory length, non-stationarity, and time horizon.
Study shows high costs for replicating financial claims with fixed fees.
problem High costs for replicating financial claims in markets with fixed transaction costs.
method Stochastic impulse control problem with terminal state constraint.
result Super--replication prices are prohibitively costly and lead to trivial strategies in continuous models.
We revisit the optimal investment and consumption model of Davis and Norman (1990) and Shreve and Soner (1994), following a shadow-price approach similar to that of Kallsen and Muhle-Karbe (2010). Making use of the completeness of the model without transaction costs, we reformulate and reduce the Hamilton-Jacobi-Bellma…
The paper improves competitive and dynamic regret bounds for smoothed online learning.
problem Smoothed online learning with hitting and switching costs.
method Optimization problems to minimize hitting cost, dynamic regret modification of existing algorithms.
result Improved competitive and dynamic regret bounds for various function classes.
A new cost-frugal HPO method controls training cost during optimization.
problem Ignoring training cost variation in HPO leads to inefficient hyperparameter tuning.
method Developed a randomized direct-search method with convergence and approximation guarantees.
result Proved an O ( d K ) O(\frac{\sqrt{d}}{\sqrt{K}}) O ( K d ) convergence rate and O ( d ε − 2 ) O(dε^{-2}) O ( d ε − 2 ) approximation guarantee. Incomplete financial markets are considered, defined by a multi-dimensional non-homogeneous diffusion process, being the direct sum of an Itô process (the price process), and another non-homogeneous diffusion process (the exogenous process, representing exogenous stochastic sources). The drift and the diffusion matrix …
Model cash management under ambiguity using maxmin preferences and diffusion.
problem Optimizing cash reserves in the presence of ambiguity.
method Singular control model with maxmin preferences, verified using Dynkin games.
result Higher expected costs and narrower inaction region under increased ambiguity.
The paper teaches robots to navigate by learning costs from expert demonstrations.
problem Teaching robots to navigate autonomously using only expert observations.
method Developed a map encoder and cost encoder to infer semantic class probabilities and a cost function from expert observations.
result Robots can learn to follow traffic rules in a simulator using only semantic observations.