This paper addresses the problem of learning the optimal control policy for a nonlinear stochastic dynamical system with continuous state space, continuous action space and unknown dynamics. This class of problems are typically addressed in stochastic adaptive control and reinforcement learning literature using model-b…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DQN outperforms static policies in a dynamic fee environment for automated market makers.
Framework learns robust control policies from expert demonstrations.
We propose a reinforcement learning (RL) based closed loop power control algorithm for the downlink of the voice over LTE (VoLTE) radio bearer for an indoor environment served by small cells. The main contributions of our paper are to 1) use RL to solve performance tuning problems in an indoor cellular network for voic…
Fault detection problem for closed loop uncertain dynamical systems, is investigated in this paper, using different deep learning based methods. Traditional classifier based method does not perform well, because of the inherent difficulty of detecting system level faults for closed loop dynamical system. Specifically, …
Combines Gaussian processes and polynomial chaos for stochastic control.
Study bounds the length of shortest periodic geodesics on certain curved spaces.
ApolloRL offers a platform for RL research in autonomous driving.
For large-scale industrial processes under closed-loop control, process dynamics directly resulting from control action are typical characteristics and may show different behaviors between real faults and normal changes of operating conditions. However, conventional distributed monitoring approaches do not consider the…
Paper predicts recycling bin full events to reduce RVM downtime.
Study on parameter dynamics in exponential families under closed-loop learning.
Geometrically characterizes virtual nonlinear nonholonomic constraints using symplectic methods.
Closed loop solitons in a plane, whose curvatures obey the modified Korteweg-de Vries equation, were investigated. It was shown that their tangential vectors are expressed by ratio of Weierstrass sigma functions for genus one case and ratio of Baker's sigma functions for the genus two case. This study is closely relate…
Formula adjusts steady-state models for control confounding.
In this work, we propose a new learning approach for autonomous navigation and landing of an Unmanned-Aerial-Vehicle (UAV). We develop a multimodal fusion of deep neural architectures for visual-inertial odometry. We train the model in an end-to-end fashion to estimate the current vehicle pose from streams of visual an…
The paper provides guarantees for feedback control with sensor errors.
Data-efficient reinforcement learning (RL) in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems. We consider a particularly important instance of this challenge, the pixels-to-torques problem, where an RL agent learns a closed-loop con…
Data-efficient learning in continuous state-action spaces using very high-dimensional observations remains a key challenge in developing fully autonomous systems. In this paper, we consider one instance of this challenge, the pixels to torques problem, where an agent must learn a closed-loop control policy from pixel i…
Neural Networks (NN) have been proposed in the past as an effective means for both modeling and control of systems with very complex dynamics. However, despite the extensive research, NN-based controllers have not been adopted by the industry for safety critical systems. The primary reason is that systems with learning…
Deep neural networks are known to be fragile to small adversarial perturbations. This issue becomes more critical when a neural network is interconnected with a physical system in a closed loop. In this paper, we show how to combine recent works on neural network certification tools (which are mainly used in static set…
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and fr…
Study shows how multiple traders can trade together without excessive price impact.
Modern treatments for Type 1 diabetes (T1D) use devices known as artificial pancreata (APs), which combine an insulin pump with a continuous glucose monitor (CGM) operating in a closed-loop manner to control blood glucose levels. In practice, poor performance of APs (frequent hyper- or hypoglycemic events) is common en…
Robot science discovers new materials faster.
This work discusses a closed-loop control strategy for complex systems utilizing scarce and streaming data. A discrete embedding space is first built using hash functions applied to the sensor measurements from which a Markov process model is derived, approximating the complex system's dynamics. A control strategy is t…
This paper uses NLDT to find interpretable control rules from complex DRL policies.
CLQT benchmarks LLM portfolio managers by evaluating their decision-making process, not just returns.
End-to-end learnable network for safer self-driving with interpretable intermediate representations.
Analyzes how learning algorithms affect and are affected by data manipulation.
Traditional collaborative filtering (CF) based recommender systems tend to perform poorly when the user-item interactions/ratings are highly scarce. To address this, we propose a learning framework that improves collaborative filtering with a synthetic feedback loop (CF-SFL) to simulate the user feedback. The proposed …
A homothety surface can be assembled from polygons by identifying their edges in pairs via homotheties, which are compositions of translation and scaling. We consider linear trajectories on a 1-parameter family of genus-2 homothety surfaces. The closure of a trajectory on each of these surfaces always has Hausdorff dim…
Motivated by vision-based control of autonomous vehicles, we consider the problem of controlling a known linear dynamical system for which partial state information, such as vehicle position, is extracted from complex and nonlinear data, such as a camera image. Our approach is to use a learned perception map that predi…
Deep RL improves blood glucose control for T1D patients.
Study finds loops with specific curvature exist using Hardy's inequality.
MPC outperforms reactive budgeting in non-stationary return environments.
Paper tackles stochastic control with mean and higher-order moments, finding Nash equilibria.
ANFIS system improves satellite attitude estimation and control.
A geometric approach to differential game theory is illustrated. The parallel pursuit is considered as a two-player zero-sum differential game. The optimal strategies of each player is designed based on Riemann-Finsler geometry. Our approach incorporates a closed loop optimal control and the presentation is familiar wi…
Researchers develop a method to control nonlinear systems with Koopman operator regression.
New budget quantifies drift in closed-loop learning, improving reproducibility.
Model-free Reinforcement Learning (RL) works well when experience can be collected cheaply and model-based RL is effective when system dynamics can be modeled accurately. However, both assumptions can be violated in real world problems such as robotics, where querying the system can be expensive and real-world dynamics…
We present an alternative local definition of the writhe of a self-avoiding closed loop which differs from the traditional non-local definition by an integer. When studying dynamics this difference is immaterial. We employ a formula due to Aldinger, Klapper and Tabor for the change in writhe and propose a set of local,…
PenduMAV is a 6-input omnidirectional MAV without internal forces.
The paper tackles performative risk optimization under weak convexity assumptions.
Deep RL controls anesthesia more accurately than traditional methods.
Filling length measures the length of the contracting closed loops in a null-homotopy. The filling length function of Gromov for a finitely presented group measures the filling length as a function of length of edge-loops in the Cayley 2-complex. We give a bound on the filling length function in terms of the log of an …
The feasibility of existing data stream algorithms is often hindered by the weakly supervised condition of data streams. A self-evolving deep neural network, namely Parsimonious Network (ParsNet), is proposed as a solution to various weakly-supervised data stream problems. A self-labelling strategy with hedge (SLASH) i…
Sequential learning of tasks using gradient descent leads to an unremitting decline in the accuracy of tasks for which training data is no longer available, termed catastrophic forgetting. Generative models have been explored as a means to approximate the distribution of old tasks and bypass storage of real data. Here …