Meta-learning control algorithm with finite-time guarantees for unknown systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Iterative method learns unknown constraints for MPC control.
New algorithm minimizes worst-case regret in uncertain, time-varying dynamics.
This work uses QPGPs to improve ILC performance in repetitive tasks.
Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is analyzed in solving the problem. The convergence of the iterations and the optim…
Improving optimization for iterate-averaged language models
PACE optimizes training for averaged language models, improving performance.
Adaptive optimal control using value iteration initiated from a stabilizing control policy is theoretically analyzed in terms of stability of the system during the learning stage without ignoring the effects of approximation errors. This analysis includes the system operated using any single/constant resulting control …
MVPI framework optimizes risk in reinforcement learning, improving performance in robot simulations.
We consider in this paper the regularity problem for time-optimal trajectories of a single-input control-affine system on a n-dimensional manifold. We prove that, under generic conditions on the drift and the controlled vector field, any control u associated with an optimal trajectory is smooth out of a countable set o…
Generalising the idea of the classical EM algorithm that is widely used for computing maximum likelihood estimates, we propose an EM-Control (EM-C) algorithm for solving multi-period finite time horizon stochastic control problems. The new algorithm sequentially updates the control policies in each time period using Mo…
This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of the Hamilton-Jacobi-Bellman (HJB) equation, which is a nonlinear partial differ…
This study is aimed at answering the famous question of how the approximation errors at each iteration of Approximate Dynamic Programming (ADP) affect the quality of the final results considering the fact that errors at each iteration affect the next iteration. To this goal, convergence of Value Iteration scheme of ADP…
Robust -learning for mean-field control under Wasserstein uncertainty
This paper presents several numerical applications of deep learning-based algorithms that have been introduced in [HPBL18]. Numerical and comparative tests using TensorFlow illustrate the performance of our different algorithms, namely control learning by performance iteration (algorithms NNcontPI and ClassifPI), contr…
This paper provides a full controlled version of algebraic -theory. This includes a rich array of assembly maps; the controlled assembly isomorphism theorem identifying the controlled group with homology; and the stability theorem describing the behavior of the inverse limit as the control parameter goes to 0. There…
Unified reinforcement learning methods using hybrid inference.
In this paper, we present a novel penalty approach for the numerical solution of continuously controlled HJB equations and HJB obstacle problems. Our results include estimates of the penalisation error for a class of penalty terms, and we show that variations of Newton's method can be used to obtain globally convergent…
Develops a reinforcement learning algorithm for learning deterministic equilibrium policies in time-inconsistent control problems.
Gaussian Processes (GPs) are widely employed in control and learning because of their principled treatment of uncertainty. However, tracking uncertainty for iterative, multi-step predictions in general leads to an analytically intractable problem. While approximation methods exist, they do not come with guarantees, mak…
This work is motivated by numerical solutions to Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) associated with combined stochastic and impulse control problems. In particular, we consider (i) direct control, (ii) penalized, and (iii) semi-Lagrangian discretization schemes applied to the HJBQVI proble…
New framework calibrates models to control risk under performativity.
New geometric approach controls motion of a spinning sphere on a plane.
Optimizes CM for stochastic convex optimization with progressive precision.
Study efficient iterative method for distribution matching using sliced optimal transport.
In this paper, two Q-learning (QL) methods are proposed and their convergence theories are established for addressing the model-free optimal control problem of general nonlinear continuous-time systems. By introducing the Q-function for continuous-time systems, policy iteration based QL (PIQL) and value iteration based…
We prove that the control polygon of a Bezier curve B becomes homeomorphic and ambient isotopic to B via subdivision, and we provide closed-form formulas to compute the number of iterations to ensure these topological characteristics. We first show that the exterior angles of control polygons converge exponentially to …
Model-free approaches for reinforcement learning (RL) and continuous control find policies based only on past states and rewards, without fitting a model of the system dynamics. They are appealing as they are general purpose and easy to implement; however, they also come with fewer theoretical guarantees than model-bas…
New algorithm achieves logarithmic regret for adversarial online control.
Investment strategy optimizes risk using a specific risk measure.
The paper reclassifies RL algorithms using inference concepts.
This paper analyzes a simplified strategy for nonlinear control using local linear models and iLQR updates.
We propose a method for efficient training of Q-functions for continuous-state Markov Decision Processes (MDPs) such that the traces of the resulting policies satisfy a given Linear Temporal Logic (LTL) property. LTL, a modal logic, can express a wide range of time-dependent logical properties (including "safety") that…
We analyze stochastic approximation with Markov noise for reinforcement learning.
Framework for controlling multiple risks in AI models.
We propose a novel reformulation of the stochastic optimal control problem as an approximate inference problem, demonstrating, that such a interpretation leads to new practical methods for the original problem. In particular we characterise a novel class of iterative solutions to the stochastic optimal control problem …
Adaptive optimal control using value iteration (VI) initiated from a stabilizing policy is theoretically analyzed in various aspects including the continuity of the result, the stability of the system operated using any single/constant resulting control policy, the stability of the system operated using the evolving/ti…
Market makers optimize trading with a new implicit scheme for complex inequalities.
The paper proves index theorems for graph-based optimal control problems.
Deep learning solves complex stochastic control with jumps.
Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alternately optimizes, two policies: a fast, reactive policy (e.g., a deep neural network) deployed at t…
This work studies the contraction coefficients of Schrödinger bridge problems in linear systems.
New approach uses hinge loss for iterative regularization in classification.
The paper solves complex control problems using neural networks.
Neural ODEs simplified using Chen-Fliess series for Rademacher complexity analysis.
Efficiently compress pretrained models using RSI for improved predictive accuracy.
Optimizes control of noisy discrete systems without system matrix knowledge.
Combines Lyapunov functions with controller synthesis for safe control policies.