Data-efficient reinforcement learning (RL) in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems. We consider a particularly important instance of this challenge, the pixels-to-torques problem, where an RL agent learns a closed-loop con…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Data-efficient learning in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems. In this paper, we consider one instance of this challenge, the pixels to torques problem, where an agent must learn a closed-loop control policy from pixel i…
Framework learns robust control policies from expert demonstrations.
The paper tackles performative risk optimization under weak convexity assumptions.
We utilize Wi-Fi communications from smartphones to predict their mobility mode, i.e. walking, biking and driving. Wi-Fi sensors were deployed at four strategic locations in a closed loop on streets in downtown Toronto. Deep neural network (Multilayer Perceptron) along with three decision tree based classifiers (Decisi…
Fault detection problem for closed loop uncertain dynamical systems, is investigated in this paper, using different deep learning based methods. Traditional classifier based method does not perform well, because of the inherent difficulty of detecting system level faults for closed loop dynamical system. Specifically, …
Combines Gaussian processes and polynomial chaos for stochastic control.
Paper predicts recycling bin full events to reduce RVM downtime.
Study bounds the length of shortest periodic geodesics on certain curved spaces.
For large-scale industrial processes under closed-loop control, process dynamics directly resulting from control action are typical characteristics and may show different behaviors between real faults and normal changes of operating conditions. However, conventional distributed monitoring approaches do not consider the…
Study on parameter dynamics in exponential families under closed-loop learning.
Analyzes how learning algorithms affect and are affected by data manipulation.
Bubblewrap predicts neural dynamics online, scaling to thousands of neurons.
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and fr…
Geometrically characterizes virtual nonlinear nonholonomic constraints using symplectic methods.
Closed loop solitons in a plane, whose curvatures obey the modified Korteweg-de Vries equation, were investigated. It was shown that their tangential vectors are expressed by ratio of Weierstrass sigma functions for genus one case and ratio of Baker's sigma functions for the genus two case. This study is closely relate…
This paper addresses the problem of learning the optimal control policy for a nonlinear stochastic dynamical system with continuous state space, continuous action space and unknown dynamics. This class of problems are typically addressed in stochastic adaptive control and reinforcement learning literature using model-b…
Formula adjusts steady-state models for control confounding.
New budget quantifies drift in closed-loop learning, improving reproducibility.
MPC outperforms reactive budgeting in non-stationary return environments.
End-to-end learnable network for safer self-driving with interpretable intermediate representations.
A new approach optimizes weights in DLP for better risk-adjusted performance.
Deep neural networks are known to be fragile to small adversarial perturbations. This issue becomes more critical when a neural network is interconnected with a physical system in a closed loop. In this paper, we show how to combine recent works on neural network certification tools (which are mainly used in static set…
Study shows how multiple traders can trade together without excessive price impact.
Geometrical and appearance quality requirements set the limits of the current industrial performance in injection molding. To guarantee the product's quality, it is necessary to adjust the process settings in a closed loop. Those adjustments cannot rely on the final quality because a part takes days to be geometrically…
Robot science discovers new materials faster.
PandaAI: A practical agent for neuro-symbolic data analysis and decision-making in finance
This work discusses a closed-loop control strategy for complex systems utilizing scarce and streaming data. A discrete embedding space is first built using hash functions applied to the sensor measurements from which a Markov process model is derived, approximating the complex system's dynamics. A control strategy is t…
Proposes a recursive MPC scheme with probabilistic safety guarantees for uncertain dynamic systems.
Researchers develop a method to control nonlinear systems with Koopman operator regression.
This paper uses NLDT to find interpretable control rules from complex DRL policies.
New framework optimizes forecasting and decision-making in dynamic systems.
CLQT benchmarks LLM portfolio managers by evaluating their decision-making process, not just returns.
Deep RL improves blood glucose control for T1D patients.
We propose a reinforcement learning (RL) based closed loop power control algorithm for the downlink of the voice over LTE (VoLTE) radio bearer for an indoor environment served by small cells. The main contributions of our paper are to 1) use RL to solve performance tuning problems in an indoor cellular network for voic…
TIE framework detects out-of-distribution samples and estimates uncertainty without external datasets.
DISTANA predicts and denoises spatial wave dynamics.
DQN outperforms static policies in a dynamic fee environment for automated market makers.
A homothety surface can be assembled from polygons by identifying their edges in pairs via homotheties, which are compositions of translation and scaling. We consider linear trajectories on a 1-parameter family of genus-2 homothety surfaces. The closure of a trajectory on each of these surfaces always has Hausdorff dim…
In this work, we propose a new learning approach for autonomous navigation and landing of an Unmanned-Aerial-Vehicle (UAV). We develop a multimodal fusion of deep neural architectures for visual-inertial odometry. We train the model in an end-to-end fashion to estimate the current vehicle pose from streams of visual an…
Study finds loops with specific curvature exist using Hardy's inequality.
Paper tackles stochastic control with mean and higher-order moments, finding Nash equilibria.
A geometric approach to differential game theory is illustrated. The parallel pursuit is considered as a two-player zero-sum differential game. The optimal strategies of each player is designed based on Riemann-Finsler geometry. Our approach incorporates a closed loop optimal control and the presentation is familiar wi…
Model-free Reinforcement Learning (RL) works well when experience can be collected cheaply and model-based RL is effective when system dynamics can be modeled accurately. However, both assumptions can be violated in real world problems such as robotics, where querying the system can be expensive and real-world dynamics…
In this work, we take a representation learning perspective on hierarchical reinforcement learning, where the problem of learning lower layers in a hierarchy is transformed into the problem of learning trajectory-level generative models. We show that we can learn continuous latent representations of trajectories, which…
Paper identifies tensor ranks via prior predictive matching, solving system of equations.
Motivated by vision-based control of autonomous vehicles, we consider the problem of controlling a known linear dynamical system for which partial state information, such as vehicle position, is extracted from complex and nonlinear data, such as a camera image. Our approach is to use a learned perception map that predi…
We present an alternative local definition of the writhe of a self-avoiding closed loop which differs from the traditional non-local definition by an integer. When studying dynamics this difference is immaterial. We employ a formula due to Aldinger, Klapper and Tabor for the change in writhe and propose a set of local,…