We consider the problem of online planning in a Markov Decision Process when given only access to a generative model, restricted to open-loop policies - i.e. sequences of actions - and under budget constraint. In this setting, the Open-Loop Optimistic Planning (OLOP) algorithm enjoys good theoretical guarantees but is …
Advocates a local feedback approach for RL in unknown systems.
problem Finding optimal feedback laws in unknown nonlinear dynamical systems.
method Searches over a local feedback representation consisting of an open-loop sequence and an optimal linear feedback law.
result Results in highly efficient training and superior performance compared to global methods.
Driven by the need for parallelizable hyperparameter optimization methods, this paper studies \emph{open loop} search methods: sequences that are predetermined and can be generated before a single configuration is evaluated. Examples include grid search, uniform random search, low discrepancy sequences, and other sampl…
Paper uses deep reinforcement learning for better control of rocket engines during start-up phases.
problem Lack of optimal control during transient phases of liquid rocket engines.
method Deep reinforcement learning approach for optimal control of a gas-generator engine's continuous start-up phase.
result Deep reinforcement learning controller achieves highest performance and minimal computational effort.
In this paper, we study a time-inconsistent consumption-investment problem with random endowments in a possibly incomplete market under general discount functions. We provide a necessary condition and a verification theorem for an open-loop equilibrium consumption-investment pair in terms of a coupled forward-backward …
This paper addresses the problem of learning the optimal control policy for a nonlinear stochastic dynamical system with continuous state space, continuous action space and unknown dynamics. This class of problems are typically addressed in stochastic adaptive control and reinforcement learning literature using model-b…
Decouples critic chunk length from policy to improve policy reactivity and performance.
problem Bootstrapping bias and difficulty in extracting optimal policies from chunked critics.
method Optimizes policy against a distilled critic for partial action chunks, allowing shorter chunks for policy.
result Reliably outperforms prior methods on long-horizon offline goal-conditioned tasks.
Method improves volatility targeting for index construction.
problem High turnover, leverage spikes, and sensitivity to estimation error in existing volatility-targeting strategies.
method Proportional-control approach for setting index weights that corrects tracking error through feedback.
result The proportional-control approach achieves the target volatility more effectively than open-loop alternatives.
Study on estimating unstable open-loop matrices from state trajectories.
problem System identification for stochastic continuous-time dynamics.
method Employing randomized control inputs to estimate unstable open-loop matrix.
result Estimation error decays with trajectory length, signal-to-noise ratio, and excitability.
In the context of tree-search stochastic planning algorithms where a generative model is available, we consider on-line planning algorithms building trees in order to recommend an action. We investigate the question of avoiding re-planning in subsequent decision steps by directly using sub-trees as action recommender. …
Study time-inconsistent portfolio optimization for competitive agents with relative performance criteria.
problem Time-inconsistent mean field and n-agent games under relative performance criteria.
method Construct open-loop equilibrium strategies for n-agent games and mean field games.
result Explicit solutions for n-agent games and mean field games, unique in a special class of equilibria.
Study analyzes portfolio liquidation games influenced by self-exciting order flow.
problem Analyzing portfolio liquidation strategies with market order dynamics.
method Mean-field control problem, novel FBSDE system, sufficient maximum principle.
result Existence and uniqueness of open-loop Nash equilibria proved.
Paper tackles stochastic control with mean and higher-order moments, finding Nash equilibria.
problem Time-inconsistent stochastic control problems with mean and higher-order moments.
method Developed closed-loop and open-loop Nash equilibrium controls using PDEs and maximum principles.
result Identical closed-loop and open-loop Nash equilibria controls, independent of state value and random path.
In many cases an intelligent agent may want to learn how to mimic a single observed demonstrated trajectory. In this work we consider how to perform such procedural learning from observation, which could help to enable agents to better use the enormous set of video data on observation sequences. Our approach exploits t…
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and fr…
Action chunking and data exploration improve behavior cloning in robotics.
problem Exponential errors in learning from demonstrations for continuous control tasks.
method Action chunking and exploratory data collection.
result Control-theoretic stability is key to improving imitation learning.
Control Contraction Metrics (CCMs) provide a nonlinear controller design involving an offline search for a Riemannian metric and an online search for a shortest path between the current and desired trajectories. In this paper, we generalize CCMs to Finsler geometry, allowing the use of non-Riemannian metrics. We provid…
Investigates time-inconsistent portfolio selection under MMV preferences.
problem Time-inconsistent optimal strategies for MMV preferences.
method Nash equilibrium controls for MMV and MV preferences, solving FBSDE and HJB equations.
result MMV optimal strategies lead to higher investment amounts than MV strategies, narrowing over time.
A new framework reduces inconsistencies in chaotic surrogate modeling.
problem Consistency issues between probabilistic objectives and dynamical system dynamics.
method KAFFEE (Kalman-Aware Framework For Ergodic Emulation), a differentiable extended Kalman filter.
result KAFFEE mitigates the dynamic-probabilistic consistency gap, improving reconstruction and predictive scores.
Paper proposes a deep learning method for better IMU gyroscope data.
problem Improving accuracy of IMU gyroscope data for robot orientation estimation.
method Dilated convolution neural network, proper loss function, key points identification.
result Algorithm outperforms state-of-the-art on unseen test sequences.
Develops a method to plan exploration that learns strong policies with fewer samples.
problem Lack of efficient exploration in reinforcement learning for real-world tasks.
method Plans an action sequence that maximizes information gain about the optimal trajectory.
result 2x fewer samples than exploration baselines and 200x fewer than model-free methods.
New framework optimizes forecasting and decision-making in dynamic systems.
problem Optimizing forecasting and decision-making processes in dynamic systems.
method Closed-loop framework using bilevel optimization.
result The proposed methodology yields consistently better performance than the standard open-loop approach.
Framework for robust control in cooperative systems with uncertain common noise.
problem Optimizing collective behavior of agents in the presence of uncertain common noise.
method Proposes a robust mean-field control framework and proves existence of optimal controls.
result Existence of optimal open-loop controls linked to a lifted robust Markov decision problem.
In this paper, we consider the asset-liability management under the mean-variance criterion. The financial market consists of a risk-free bond and a stock whose price process is modeled by a geometric Brownian motion. The liability of the investor is uncontrollable and is modeled by another geometric Brownian motion. W…
Study shows how multiple traders can trade together without excessive price impact.
problem Coordination issues in trading to exploit a common signal.
method Closed-loop Nash competition model for stochastic differential games.
result Excessive trading reduced but not significantly for practical parameters.
The paper analyzes strategic irreversible investments with novel dynamic strategies.
problem Tradeoff between preemption incentives and option value of waiting in oligopolistic markets.
method Developed novel Markov perfect equilibrium to handle singular control of optimal investment.
result Simpler strategies lead to a 'preemption trap' with zero net present values.
Improved Frank-Wolfe algorithm for generalized self-concordant functions converges quickly.
problem Efficiently solving learning problems with generalized self-concordant objectives.
method Simple Frank-Wolfe variant with open-loop step size strategy γt=2/(t+2). result Achieves O(1/t) convergence rate for primal and Frank-Wolfe gaps. We study the competition of two strategic agents for liquidity in the benchmark portfolio tracking setup of Bank, Soner, Voß (2017). Specifically, both agents track their own stochastic running trading targets while interacting through common aggregated temporary and permanent price impact à la Almgren and Chriss (2001…
We propose a model of inter-bank lending and borrowing which takes into account clearing debt obligations. The evolution of log-monetary reserves of N banks is described by coupled diffusions driven by controls with delay in their drifts. Banks are minimizing their finite-horizon objective functions which take into a…
In this paper, we formulate a general time-inconsistent stochastic linear--quadratic (LQ) control problem. The time-inconsistency arises from the presence of a quadratic term of the expected state as well as a state-dependent term in the objective functional. We define an equilibrium, instead of optimal, solution withi…
The goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challengin…
Investigates RI strategies for life insurers with LRD mortality rates.
problem Effect of long-range dependent mortality rates on RI strategies.
method Volterra mortality model, compound Poisson process, open-loop equilibrium mean-variance criterion.
result Explicit equilibrium RI controls derived and uniqueness studied.
The paper introduces new metrics for evaluating generative models of behavior.
problem Lack of quantitative evaluation criteria for unsupervised behavior discovery.
method Proposed and investigated several metrics for generative models of behavior.
result The proposed metrics correspond with biologists' intuitions and allow for model evaluation and bias understanding.
Neural operators approximate Stackelberg game solutions.
problem Intractability of follower's best-response operator in dynamic Stackelberg games.
method Used attention-based neural operators to approximate the best-response operator.
result Approximate best-response operator yields close game value.
FalconBC improves patient-specific cardiovascular modeling by estimating boundary conditions efficiently.
problem Efficiently estimating boundary conditions in patient-specific cardiovascular models, especially in open-loop models and anatomies with lesions.
method A general amortized inference framework based on probabilistic flow that treats clinical targets and anatomies as conditioning variables.
result Demonstrated on two patient-specific models, FalconBC improves efficiency and accuracy in estimating boundary conditions.
Cointegration helps insurers understand long-range mortality patterns.
problem Insurers struggle to detect long-range dependence in their mortality data.
method Cointegration techniques applied to mixed fractional Brownian motion (mfBm) to capture long-range dependence.
result Cointegration brings long-range dependence information from national mortality data to insurers' models.
Improved neural ODEs learn adaptable flows.
problem Neural ODEs struggle with expressive power and adaptability.
method Introduce N-CODE modules with dynamic parameters controlled by a trainable map.
result N-CODE modules enhance expressivity of neural ODEs.
The design of multiple experiments is commonly undertaken via suboptimal strategies, such as batch (open-loop) design that omits feedback or greedy (myopic) design that does not account for future effects. This paper introduces new strategies for the optimal design of sequential experiments. First, we rigorously formul…
Model explains periodic trading in financial markets through game theory.
problem Understanding periodic trading activities in financial markets.
method Mean-field liquidation game with major-minor players.
result Existence and uniqueness of Nash equilibrium established.
In this paper, we apply the idea of fictitious play to design deep neural networks (DNNs), and develop deep learning theory and algorithms for computing the Nash equilibrium of asymmetric N-player non-zero-sum stochastic differential games, for which we refer as \emph{deep fictitious play}, a multi-stage learning pro…
We show LLMs can be locally linear, enabling better control of activations.
problem Suboptimal control of LLM activations during generation.
method Model LLM inference as a linear dynamical system, compute feedback controllers using Jacobians, and adapt classical control theory.
result Robust, fine-grained control of LLM activations across models and tasks.
Study non-parametric frequency-domain system identification from finite samples.
problem Frequency-domain system identification from limited data.
method Empirical Transfer Function Estimate (ETFE) under sub-Gaussian colored noise and stability assumptions.
result ETFE estimates are concentrated around true values with a finite-sample rate of Ntot−1/3 for all frequencies in the H∞ norm. Paper characterizes equilibrium strategies for stochastic control with higher-order moments.
problem Stochastic control problems with higher-order moments.
method Novel characterization of time-consistent control problems, deriving equilibrium conditions via BSDEs.
result Derives sufficient and necessary conditions for an open-loop Nash equilibrium control (ONEC) in a novel way.
We study the out-of-sample properties of robust empirical optimization problems with smooth φ-divergence penalties and smooth concave objective functions, and develop a theory for data-driven calibration of the non-negative "robustness parameter" δ that controls the size of the deviations from the nominal model. Bu…
Stablecoins offer efficient settlement but externalize costs and risks.
problem Comparing stablecoins to card networks in retail payments.
method Unified analytical framework (CLEAR) across five dimensions.
result Stablecoins are advantageous in closed-loop and high-friction contexts but structurally disadvantaged as open-loop instruments.
Neural models learn continuous-time Markov chain transition rates from data.
problem Learning transition rates for complex stochastic systems.
method Neural networks to model nonlinear transition rates from observed data.
result Neural models outperform traditional methods in accuracy.
A novel sequence-to-sequence model predicts missing sensor data.
problem Missing sensor data in sequences.
method Formulated a novel sequence-to-sequence model using forward and backward RNNs.
result The model produces the lowest errors in 12% more cases than the current state-of-the-art.
Combining causality, control, and reinforcement learning for system control.
problem Learning to control dynamical systems using causal, control, and reinforcement learning approaches.
method Combining causal identification, control strategies, and reinforcement learning to control dynamical systems.
result Combining different learning paradigms for effective system control.