Improved algorithm for online planning with tighter bounds.
problem Online planning in Markov Decision Processes with open-loop policies and budget constraints.
method Proposed KLOLOP algorithm with tighter upper-confidence bounds and efficient implementation.
result KLOLOP leads to better practical performances with sample complexity bound retained.
Novel algorithm for optimal control of nonlinear systems.
problem Optimal control of nonlinear stochastic dynamical systems with unknown dynamics.
method Decoupled data-based approach combining open-loop and closed-loop control.
result Performance of D2C algorithm is approximately optimal and significantly reduces training time.
Decouples critic chunk length from policy to improve policy reactivity and performance.
problem Bootstrapping bias and difficulty in extracting optimal policies from chunked critics.
method Optimizes policy against a distilled critic for partial action chunks, allowing shorter chunks for policy.
result Reliably outperforms prior methods on long-horizon offline goal-conditioned tasks.
Paper uses deep reinforcement learning for better control of rocket engines during start-up phases.
problem Lack of optimal control during transient phases of liquid rocket engines.
method Deep reinforcement learning approach for optimal control of a gas-generator engine's continuous start-up phase.
result Deep reinforcement learning controller achieves highest performance and minimal computational effort.
Paper proposes a method for optimizing local policies for trajectory-centric reinforcement learning.
problem Challenges in global policy optimization for non-linear systems and poor performance of open-loop trajectory optimization.
method Formulates trajectory optimization and local policy synthesis as a single optimization problem and solves it as a nonlinear programming instance.
result Demonstrates improved performance of the proposed technique under simplifying assumptions.
Develops a method to plan exploration that learns strong policies with fewer samples.
problem Lack of efficient exploration in reinforcement learning for real-world tasks.
method Plans an action sequence that maximizes information gain about the optimal trajectory.
result 2x fewer samples than exploration baselines and 200x fewer than model-free methods.
Advocates a local feedback approach for RL in unknown systems.
problem Finding optimal feedback laws in unknown nonlinear dynamical systems.
method Searches over a local feedback representation consisting of an open-loop sequence and an optimal linear feedback law.
result Results in highly efficient training and superior performance compared to global methods.
PLOTS learns procedural actions from observed sequences, up to 100x faster.
problem Learning procedural actions from observed sequences efficiently.
method Exploits subtask structure to incrementally build action plans, optimistically explores actions.
result Explicit procedural learning is 100x faster than policy-gradient methods and model-based approaches.
Study time-inconsistent consumption-investment in incomplete markets with general discount functions.
problem Time-inconsistent consumption-investment problems in incomplete markets.
method Coupled forward-backward stochastic differential equation approach.
result Uniqueness of open-loop equilibrium pair proved.
The design of multiple experiments is commonly undertaken via suboptimal strategies, such as batch (open-loop) design that omits feedback or greedy (myopic) design that does not account for future effects. This paper introduces new strategies for the optimal design of sequential experiments. First, we rigorously formul…
Method improves volatility targeting for index construction.
problem High turnover, leverage spikes, and sensitivity to estimation error in existing volatility-targeting strategies.
method Proportional-control approach for setting index weights that corrects tracking error through feedback.
result The proportional-control approach achieves the target volatility more effectively than open-loop alternatives.
Study on estimating unstable open-loop matrices from state trajectories.
problem System identification for stochastic continuous-time dynamics.
method Employing randomized control inputs to estimate unstable open-loop matrix.
result Estimation error decays with trajectory length, signal-to-noise ratio, and excitability.
This paper uses NLDT to find interpretable control rules from complex DRL policies.
problem Complex, non-interpretable policies from black-box AI methods.
method Evolutionary optimization of NLDT for hierarchical control rules.
result Interpretable control rules with similar performance to black-box DRL.
In the context of tree-search stochastic planning algorithms where a generative model is available, we consider on-line planning algorithms building trees in order to recommend an action. We investigate the question of avoiding re-planning in subsequent decision steps by directly using sub-trees as action recommender. …
Study time-inconsistent portfolio optimization for competitive agents with relative performance criteria.
problem Time-inconsistent mean field and n-agent games under relative performance criteria.
method Construct open-loop equilibrium strategies for n-agent games and mean field games.
result Explicit solutions for n-agent games and mean field games, unique in a special class of equilibria.
Driven by the need for parallelizable hyperparameter optimization methods, this paper studies \emph{open loop} search methods: sequences that are predetermined and can be generated before a single configuration is evaluated. Examples include grid search, uniform random search, low discrepancy sequences, and other sampl…
Study analyzes portfolio liquidation games influenced by self-exciting order flow.
problem Analyzing portfolio liquidation strategies with market order dynamics.
method Mean-field control problem, novel FBSDE system, sufficient maximum principle.
result Existence and uniqueness of open-loop Nash equilibria proved.
Paper tackles stochastic control with mean and higher-order moments, finding Nash equilibria.
problem Time-inconsistent stochastic control problems with mean and higher-order moments.
method Developed closed-loop and open-loop Nash equilibrium controls using PDEs and maximum principles.
result Identical closed-loop and open-loop Nash equilibria controls, independent of state value and random path.
Generative model learns to mimic time-series behavior to generate accurate trajectories.
problem Learning generative models for time-series data with accurate multi-step trajectories.
method Contrastive imitation framework combining autoregressive and adversarial elements.
result The method generates accurate and useful samples from real-world datasets.
Efficient learning-based MPC for unknown nonlinear systems with state constraints.
problem Control of discrete-time nonlinear systems with unknown dynamics and state constraints.
method Receding horizon reinforcement learning (r-LPC) using Koopman operator-based prediction model.
result Proven closed-loop recursive feasibility, robustness, and asymptotic stability under function approximation errors.
Control Contraction Metrics (CCMs) provide a nonlinear controller design involving an offline search for a Riemannian metric and an online search for a shortest path between the current and desired trajectories. In this paper, we generalize CCMs to Finsler geometry, allowing the use of non-Riemannian metrics. We provid…
Investigates time-inconsistent portfolio selection under MMV preferences.
problem Time-inconsistent optimal strategies for MMV preferences.
method Nash equilibrium controls for MMV and MV preferences, solving FBSDE and HJB equations.
result MMV optimal strategies lead to higher investment amounts than MV strategies, narrowing over time.
The paper analyzes optimal overbetting strategies for a satellite investment account.
problem Optimal control of leverage in a satellite investment account with limited leverage.
method Recursive overbetting strategy to maximize growth rate, solved via HJB equation.
result Optimal overbetting strategy balances growth rate of satellite and composite bankroll.
A new framework reduces inconsistencies in chaotic surrogate modeling.
problem Consistency issues between probabilistic objectives and dynamical system dynamics.
method KAFFEE (Kalman-Aware Framework For Ergodic Emulation), a differentiable extended Kalman filter.
result KAFFEE mitigates the dynamic-probabilistic consistency gap, improving reconstruction and predictive scores.
New framework optimizes forecasting and decision-making in dynamic systems.
problem Optimizing forecasting and decision-making processes in dynamic systems.
method Closed-loop framework using bilevel optimization.
result The proposed methodology yields consistently better performance than the standard open-loop approach.
Framework for robust control in cooperative systems with uncertain common noise.
problem Optimizing collective behavior of agents in the presence of uncertain common noise.
method Proposes a robust mean-field control framework and proves existence of optimal controls.
result Existence of optimal open-loop controls linked to a lifted robust Markov decision problem.
In this paper, we consider the asset-liability management under the mean-variance criterion. The financial market consists of a risk-free bond and a stock whose price process is modeled by a geometric Brownian motion. The liability of the investor is uncontrollable and is modeled by another geometric Brownian motion. W…
Study shows how multiple traders can trade together without excessive price impact.
problem Coordination issues in trading to exploit a common signal.
method Closed-loop Nash competition model for stochastic differential games.
result Excessive trading reduced but not significantly for practical parameters.
The paper analyzes strategic irreversible investments with novel dynamic strategies.
problem Tradeoff between preemption incentives and option value of waiting in oligopolistic markets.
method Developed novel Markov perfect equilibrium to handle singular control of optimal investment.
result Simpler strategies lead to a 'preemption trap' with zero net present values.
Improved Frank-Wolfe algorithm for generalized self-concordant functions converges quickly.
problem Efficiently solving learning problems with generalized self-concordant objectives.
method Simple Frank-Wolfe variant with open-loop step size strategy γt=2/(t+2). result Achieves O(1/t) convergence rate for primal and Frank-Wolfe gaps. We propose a model of inter-bank lending and borrowing which takes into account clearing debt obligations. The evolution of log-monetary reserves of N banks is described by coupled diffusions driven by controls with delay in their drifts. Banks are minimizing their finite-horizon objective functions which take into a…
In this paper, we formulate a general time-inconsistent stochastic linear--quadratic (LQ) control problem. The time-inconsistency arises from the presence of a quadratic term of the expected state as well as a state-dependent term in the objective functional. We define an equilibrium, instead of optimal, solution withi…
Investigates RI strategies for life insurers with LRD mortality rates.
problem Effect of long-range dependent mortality rates on RI strategies.
method Volterra mortality model, compound Poisson process, open-loop equilibrium mean-variance criterion.
result Explicit equilibrium RI controls derived and uniqueness studied.
Two strategic agents track their portfolios, influencing each other's trading targets.
problem Strategic competition in portfolio tracking with price impact.
method Stochastic linear quadratic differential game with terminal state constraints.
result Unique open-loop Nash equilibrium strategies emerge based on price impact types.
The paper introduces new metrics for evaluating generative models of behavior.
problem Lack of quantitative evaluation criteria for unsupervised behavior discovery.
method Proposed and investigated several metrics for generative models of behavior.
result The proposed metrics correspond with biologists' intuitions and allow for model evaluation and bias understanding.
Neural operators approximate Stackelberg game solutions.
problem Intractability of follower's best-response operator in dynamic Stackelberg games.
method Used attention-based neural operators to approximate the best-response operator.
result Approximate best-response operator yields close game value.
FalconBC improves patient-specific cardiovascular modeling by estimating boundary conditions efficiently.
problem Efficiently estimating boundary conditions in patient-specific cardiovascular models, especially in open-loop models and anatomies with lesions.
method A general amortized inference framework based on probabilistic flow that treats clinical targets and anatomies as conditioning variables.
result Demonstrated on two patient-specific models, FalconBC improves efficiency and accuracy in estimating boundary conditions.
Action chunking and data exploration improve behavior cloning in robotics.
problem Exponential errors in learning from demonstrations for continuous control tasks.
method Action chunking and exploratory data collection.
result Control-theoretic stability is key to improving imitation learning.
Cointegration helps insurers understand long-range mortality patterns.
problem Insurers struggle to detect long-range dependence in their mortality data.
method Cointegration techniques applied to mixed fractional Brownian motion (mfBm) to capture long-range dependence.
result Cointegration brings long-range dependence information from national mortality data to insurers' models.
Improved neural ODEs learn adaptable flows.
problem Neural ODEs struggle with expressive power and adaptability.
method Introduce N-CODE modules with dynamic parameters controlled by a trainable map.
result N-CODE modules enhance expressivity of neural ODEs.
Model explains periodic trading in financial markets through game theory.
problem Understanding periodic trading activities in financial markets.
method Mean-field liquidation game with major-minor players.
result Existence and uniqueness of Nash equilibrium established.
Paper proposes a deep learning method for better IMU gyroscope data.
problem Improving accuracy of IMU gyroscope data for robot orientation estimation.
method Dilated convolution neural network, proper loss function, key points identification.
result Algorithm outperforms state-of-the-art on unseen test sequences.
We show LLMs can be locally linear, enabling better control of activations.
problem Suboptimal control of LLM activations during generation.
method Model LLM inference as a linear dynamical system, compute feedback controllers using Jacobians, and adapt classical control theory.
result Robust, fine-grained control of LLM activations across models and tasks.
New method corrects state distribution mismatch for off-policy policy optimization.
problem Mismatch between behavior and evaluation policy state distributions.
method Off-policy policy gradient with state distribution correction.
result Significantly improved policy quality in simulations.
Study non-parametric frequency-domain system identification from finite samples.
problem Frequency-domain system identification from limited data.
method Empirical Transfer Function Estimate (ETFE) under sub-Gaussian colored noise and stability assumptions.
result ETFE estimates are concentrated around true values with a finite-sample rate of Ntot−1/3 for all frequencies in the H∞ norm. Adapts GRPO for off-policy RL, improving reward.
problem Improving training stability and efficiency in RL.
method Adapts GRPO to off-policy setting, uses clipped surrogate objectives.
result Off-policy GRPO outperforms on-policy GRPO in empirical tests.
Paper tackles efficient evaluation of natural stochastic policies in offline RL.
problem Efficiency issues in evaluating natural stochastic policies due to unknown evaluation policy.
method Derive efficiency bounds for tilting and modified treatment policies, propose nonparametric estimators.
result Proposed estimators attain efficiency bounds under lax conditions and enjoy partial double robustness.
New framework studies policy learning problems under data scarcity.
problem Learning improving policies when data is insufficient.
method Developed a mathematical framework for policy learning problems.
result Reduced policy learning problems to simpler ones in sample complexity.