Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

87174260347 · Jun 202019922001200920182026
48 results for control policy

Survey explores geometric aspects of policy optimization in control systems.

problem Understanding the geometric relationships between control design and optimization.
method Geometric perspective on policy optimization, focusing on parameterization and topology.
result Implications of policy geometry on stability and performance of local search algorithms.

Adapts model-based advice to stabilize black-box policies for nonlinear control.

problem Stabilizing machine-learned policies for nonlinear control with limited model information.
method Proposes an adaptive λλ-confident policy to combine black-box and model-based advice.
result Proves the stability of the adaptive λλ-confident policy and its competitive ratio.

PFPN uses particle filtering to improve character control in physics-based simulations.

problem Premature commitment to suboptimal actions in high-dimensional continuous control problems for articulated characters.
method Proposes a particle-based action policy using particle filtering to dynamically explore and discretize the action space.
result Demonstrates better imitation performance and robustness to external perturbations compared to Gaussian policies.

Paper proposes a new method to optimize robot body structure and control policy.

problem Optimizing robot body structure and control policy in a coupled manner.
method Revisits co-design problem as a Stackelberg game, incorporating control adaptation dynamics.
result Stackelberg PPO outperforms standard PPO in stability and performance.

New framework for policy gradient methods in continuous time reinforcement learning.

problem Addressing policy gradient methods for continuous time reinforcement learning.
method Control randomisation technique to derive policy gradient representation for various Markovian control problems.
result Demonstrated application to optimal switching problems in the energy sector.

Survey of theoretical foundations for policy optimization in control.

problem Understanding the theoretical properties of gradient-based methods in control and reinforcement learning.
method Interdisciplinary review of optimization landscape, convergence, and sample complexity for various control problems.
result Recent theoretical results on stability and robustness in learning-based control.

Researchers develop methods to protect reinforcement learning and control systems from policy poisoning attacks.

problem Attacks on reinforcement learning and control systems that manipulate learned policies.
method Unified framework for solving policy poisoning attacks, demonstrating global optimality and feasibility.
result Policy poisoning attacks are feasible and can be defended with a convex optimization approach.

Paper optimizes traffic signal control for better traffic flow.

problem Optimizing traffic signal control to reduce congestion and improve safety.
method Value-based reinforcement learning with interpretable policy functions (polynomial functions).
result Deep Regulatable Hardmax Q-learning variant reduces vehicle delay by up to 19.4%.

New method transfers robotic control policies to unknown environments using a family of policies.

problem Transfer of robotic control policies trained in simulation to real hardware is difficult due to environment differences.
method Simultaneously learns a family of policies with different behaviors; searches for the best policy based on task performance.
result Demonstrates superior performance in unknown environments compared to other methods, overcoming larger modeling errors.

The paper examines how macroeconomic control tools lost effectiveness, leading to a 'dark ages' period.

problem Loss of effectiveness of control tools in macroeconomic stabilization policy.
method Historical analysis of macroeconomic stabilization policy from 1948 to 1993.
result The overstatement of the Lucas critique and Kydland and Prescott's time-inconsistency led to a period of ineffective stabilization policy.

A new method improves stability in policy learning for continuous control tasks.

problem Stability issues in policy gradient methods when policies are close to deterministic.
method Target Distribution Learning (TDL) alternates between proposing a target distribution and training the policy network to approach it.
result TDL leads to more stable policy improvements over iterations compared to existing methods.

Trust-region methods and natural gradients are equivalent in certain policy search scenarios.

problem Improving policy search methods in continuous control tasks.
method Introducing compatible policy search (COPOS) that uses natural parameterization and compatible value function approximation to control entropy loss.
result COPOS yields state-of-the-art results in challenging tasks and reduces entropy loss.

This work ensures policy gradient methods converge to global optima for certain control problems.

problem Non-convex optimization challenges in policy gradient methods for complex control problems.
method Identifies structural properties ensuring non-convex objective functions have no suboptimal stationary points.
result Policy gradient methods converge to global optima under certain conditions, satisfying a Polyak-Lojasiewicz condition.

Developed policy gradient methods for stochastic control with exit time, outperforming traditional techniques in share repurchase pricing.

problem Optimal control with exit time in stochastic models.
method Two types of algorithms: direct policy learning and alternately learning value function and control.
result Policy gradient methods outperform PDE or neural networks in share repurchase pricing.

The paper develops a method to create robust control policies for robots using information bottlenecks.

problem Robotic control policies are sensitive to task-irrelevant state and sensor changes.
method Derives a policy gradient algorithm that creates an information bottleneck between states and task-relevant representations.
result Task-driven policies are more robust to sensor noise and environmental changes.

Efficient deep policy gradient method for continuous-time control problems.

problem Optimal control in continuous time with fine time discretization.
method Multi-scale deep policy gradient method with varying time discretization.
result Targeted efficiency in computational resources achieved through multi-scale approach.

A contextual bandit method evaluates and improves inventory control policies.

problem Evaluating and improving periodic review inventory control policies with nonstationary demand.
method Contextual bandit-based algorithm to evaluate and tweak policies.
result The method achieves favorable guarantees in both theory and practice.

The paper proposes a new method to estimate optimal policies using MCMC.

problem Estimating the optimal policy for systems with unknown dynamics and reward functions.
method Using Markov Chain Monte Carlo to generate samples from the posterior distribution of parameters conditioned on optimality.
result The method provably converges to the globally optimal stochastic policy with similar variance to policy gradient methods.

Researchers solve a complex financial control problem with explicit policies.

problem Constrained LQ control with multiplicative noise in financial risk management.
method Derived analytical control policy using state separation property and solving coupled Riccati equations.
result Explicit piece-wise affine control policy for optimal control of stochastic systems.

This paper tackles robust control of LQR systems with multiplicative noise using policy gradient methods.

problem Robustness in reinforcement learning control of complex systems with multiplicative noise.
method Policy gradient algorithms with gradient domination property for non-convex cost functions.
result Global convergence of policy gradient algorithms to the globally optimum control policy.

Optimal policy for multi-hypothesis testing with controlled sensing to minimize delay and error.

problem Minimizing delay in multi-hypothesis testing with controlled sensing.
method Designing a policy to control the delay while ensuring error probability constraint.
result Policy achieves information-theoretic lower bound on expected delay asymptotically.

Paper develops a distributed power control method for large energy harvesting networks using deep reinforcement learning.

problem Optimal power control for large energy harvesting networks with limited causal information.
method Multi-agent reinforcement learning framework to solve a mean-field game problem.
result Proposed method converges to optimal power control policies in a distributed fashion.

Paper uses imitation learning to create efficient insulin policies from MPC demonstrations.

problem Resource-constrained medical devices struggle with complex MPC optimizations and state estimation errors.
method Imitation learning of neural network policies from MPC-computed demonstrations, using Bayesian inference with Monte Carlo Dropout.
result Trained policies generalize well to different patient cohorts, outperforming traditional MPC with state estimation.

New method reduces over-pessimism in Bayesian control under parameter uncertainty.

problem Over-pessimism in Bayesian control due to misspecified priors.
method Distributionally robust Bayesian control (DRBC) with strong duality and optimization.
result Validated algorithm on synthetic and real data, reducing over-pessimism.

Risk-controlled post-processing optimizes decision policies under risk constraints.

problem Optimizing decision policies with risk constraints for better outcomes.
method Developed a post-processing algorithm that selects a threshold based on fitted fallback policy and score, leveraging tools from algorithmic stability and stochastic processes.
result The post-processed policy achieves precise expected risk control under exchangeability and meets or nearly meets risk budgets while preserving more agreement with the baseline.

Study nonparametric estimator for Markov chain transition matrices in offline setting.

problem Estimating transition matrices of finite controlled Markov chains from logged data.
method Developed sample complexity bounds and conditions for minimaxity.
result Achieving certain statistical risk requires balancing mixing properties and sample size.

Develops a new reinforcement learning framework for complex control problems.

problem Continuous-time extended mean field control with deterministic policies.
method Model-free sensitivity formula, deterministic policy gradient, local value and advantage-rate representations.
result Demonstrates efficiency, stability, and robustness in solving complex control problems.

Optimizes control of noisy discrete systems without system matrix knowledge.

problem Optimal control of discrete-time systems with additive and multiplicative noises.
method Stochastic Lyapunov and Riccati equations, model-free reinforcement learning.
result Model-free reinforcement learning algorithm converges to optimal control policy.

Framework learns robust control policies from expert demonstrations.

problem Adversarial robustness and closed-loop generalization in feedback control policies.
method Lipschitz-constrained loss minimization for certified robustness and generalization.
result Finite sample bound on policy learning error and robust closed-loop stability.

Paper formulates mutual information optimal control for discrete-time systems.

problem Optimal control of discrete-time linear systems with mutual information.
method Formulates MIOCP as an extension of MEOCP, derives optimal policy and prior, proposes alternating minimization algorithm.
result Proposes an alternating minimization algorithm for MIOCP.