The paper studies optimal transport in linear quadratic systems and derives interpolation inequalities.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present the first computationally-efficient algorithm with regret for learning in Linear Quadratic Control systems with unknown dynamics. By that, we resolve an open question of Abbasi-Yadkori and Szepesvári (2011) and Dean, Mania, Matni, Recht, and Tu (2018).
RL solves discrete LQ control with Gaussian optimal policy.
We study the problem of regret minimization in partially observable linear quadratic control systems when the model dynamics are unknown a priori. We propose ExpCommit, an explore-then-commit algorithm that learns the model Markov parameters and then follows the principle of optimism in the face of uncertainty to desig…
Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.
LqgOpt learns optimal control in unknown LQG systems with minimal regret.
We study the performance of the certainty equivalent controller on Linear Quadratic (LQ) control problems with unknown transition dynamics. We show that for both the fully and partially observed settings, the sub-optimality gap between the cost incurred by playing the certainty equivalent controller on the true system …
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
Policy gradient methods find Nash equilibrium in noisy games.
Study shows certainty equivalent policy minimizes regret in continuous-time systems.
New algorithm learns LQR with regret using Langevin dynamics and excitation.
This paper establishes the existence of a unique nonnegative continuous viscosity solution to the HJB equation associated with a Markovian linear-quadratic control problems with singular terminal state constraint and possibly unbounded cost coefficients. The existence result is based on a novel comparison principle for…
Optimizes dividends with stability for risky businesses.
Paper uses Koopman operator and Nyström method for efficient nonlinear control.
New algorithms improve blind source separation for linear-quadratic mixtures.
Develops a control framework for systemic risk under uncertainty.
Study on PG learning for LQ MFC problems with common noise, proving convergence and sample complexity.
We study the constrained linear quadratic regulator with unknown dynamics, addressing the tension between safety and exploration in data-driven control techniques. We present a framework which allows for system identification through persistent excitation, while maintaining safety by guaranteeing the satisfaction of st…
We consider a general time-inconsistent stochastic linear-quadratic differential game. The time-inconsistency arises from the presence of quadratic terms of the expected state as well as state-dependent term in the objective functionals. We define an equilibrium strategy, which is different from the classical one, and …
New theory extends LQ control to non-exponential discount scenarios.
Study optimizes resource allocation in noisy systems for better control.
Efficiently solves exploration-exploitation in LQR using Lagrangian relaxation.
Paper solves time-inconsistent control problems with BSDEs.
We study the global convergence of generative adversarial imitation learning for linear quadratic regulators, which is posed as minimax optimization. To address the challenges arising from non-convex-concave geometry, we analyze the alternating gradient algorithm and establish its Q-linear rate of convergence to a uniq…
Study solves HJB equations for time-inconsistent control problems.
We consider the problem of learning in Linear Quadratic Control systems whose transition parameters are initially unknown. Recent results in this setting have demonstrated efficient learning algorithms with regret growing with the square root of the number of decision steps. We present new efficient algorithms that ach…
Study optimizes trading in multiple assets with cross-effects.
Study policy gradient for large-agent mean-field control and game in continuous time.
This paper studies a class of continuous-time scalar-state stochastic Linear-Quadratic (LQ) optimal control problem with the linear control constraints. Applying the state separation theorem induced from its special structure, we develop the explicit solution for this class of problem. The revealed optimal control poli…
Study cost-driven state representation learning for control from partial observations.
Policy gradient methods converge for LQR problems with noisy state dynamics.
New model-free algorithm achieves similar LQR regret guarantees.
We consider adaptive control of the Linear Quadratic Regulator (LQR), where an unknown linear system is controlled subject to quadratic costs. Leveraging recent developments in the estimation of linear systems and in robust controller synthesis, we present the first provably polynomial time algorithm that provides high…
Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of dimensionality caused by the exponential growth of the state-action space with the number o…
This paper studies how gradient descent in control systems can perform well on unseen data.
In this paper, we continue our study on a general time-inconsistent stochastic linear--quadratic (LQ) control problem originally formulated in [6]. We derive a necessary and sufficient condition for equilibrium controls via a flow of forward--backward stochastic differential equations. When the state is one dimensional…
Theory is developed for linear-quadratic at infinity generating families for Legendrian knots in R^3. It is shown that the unknot with maximal Thurston--Bennequin invariant of -1 has a unique linear-quadratic at infinity generating family, up to fiber-preserving diffeomorphism and stabilization. From this, invariant ge…
The paper solves TIC LQ control problems using stochastic differential games.
Study learns state representations from observations for control, proving guarantees.
We show by counterexample that policy-gradient algorithms have no guarantees of even local convergence to Nash equilibria in continuous action and state space multi-agent settings. To do so, we analyze gradient-play in N-player general-sum linear quadratic games, a classic game setting which is recently emerging as a b…
Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence is known to be fragile. To understand the instability of actor-critic, we focus o…
We consider the exploration-exploitation tradeoff in linear quadratic (LQ) control problems, where the state dynamics is linear and the cost function is quadratic in states and controls. We analyze the regret of Thompson sampling (TS) (a.k.a. posterior-sampling for reinforcement learning) in the frequentist setting, i.…
This work establishes safe reinforcement learning for LQR with nonlinear baselines.
TSAC achieves optimal frequentist regret in adaptive control of LQRs.
Solves steering problem with continuous time, Hilbert-Schmidt cost, and matrix ODEs.
Study optimizes investment strategies in markets with contagious price jumps.
We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various settings of driving noise and reward feedback. We show that these methods provably converge to within a…
In this paper, we formulate a general time-inconsistent stochastic linear--quadratic (LQ) control problem. The time-inconsistency arises from the presence of a quadratic term of the expected state as well as a state-dependent term in the objective functional. We define an equilibrium, instead of optimal, solution withi…