Paper solves pendulum swing-up problem using RL.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
An energy based approach for stabilizing a mechanical system has offered a simple yet powerful control scheme. However, since it does not impose such strong constraints on parameter space of the controller, finding appropriate parameter values for an optimal controller is known to be hard. This paper intends to generat…
Controller seeks informative system observations to predict nonlinear dynamics.
A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, na…
We present an approach to identify concise equations from data using a shallow neural network approach. In contrast to ordinary black-box regression, this approach allows understanding functional relations and generalizing them from observed data to unseen parts of the parameter space. We show how to extend the class o…
Empowerment quantifies the influence an agent has on its environment. This is formally achieved by the maximum of the expected KL-divergence between the distribution of the successor state conditioned on a specific action and a distribution where the actions are marginalised out. This is a natural candidate for an intr…
Developing mathematical models of dynamic systems is central to many disciplines of engineering and science. Models facilitate simulations, analysis of the system's behavior, decision making and design of automatic control algorithms. Even inherently model-free control techniques such as reinforcement learning (RL) hav…
Physics-informed learning framework for pH systems and EB-PBC control.
Bayesian optimization outperforms other methods in hyperparameter tuning for reinforcement learning.
Reinforcement learning algorithms can solve dynamic decision-making and optimal control problems. With continuous-valued state and input variables, reinforcement learning algorithms must rely on function approximators to represent the value function and policy mappings. Commonly used numerical approximators, such as ne…
Reinforcement Learning methods are capable of solving complex problems, but resulting policies might perform poorly in environments that are even slightly different. In robotics especially, training and deployment conditions often vary and data collection is expensive, making retraining undesirable. Simulation training…
Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.
Proposes DLGPD model to learn dynamics from images for planning.
In this paper we show that there are applications that transform the movement of a pendulum into movements in . This can be done using Euler top system of differential equations. On the constant level surfaces, Euler top system reduces to the equation of a pendulum. Those properties are also considered in…
This paper studies the topology of the constant energy surfaces of the double spherical pendulum.
Study of sub-Riemannian problem on specific Lie groups, revealing symmetries and bounds.
This study presents the results of a series of simulation experiments that evaluate and compare four different manifold alignment methods under the influence of noise. The data was created by simulating the dynamics of two slightly different double pendulums in three-dimensional space. The method of semi-supervised fea…
This paper classifies Legendre singularities of sub-Riemannian geodesics on surfaces.
Designing optimal controllers continues to be challenging as systems are becoming complex and are inherently nonlinear. The principal advantage of reinforcement learning (RL) is its ability to learn from the interaction with the environment and provide optimal control strategy. In this paper, RL is explored in the cont…
We present a data-efficient reinforcement learning algorithm resistant to observation noise. Our method extends the highly data-efficient PILCO algorithm (Deisenroth & Rasmussen, 2011) into partially observed Markov decision processes (POMDPs) by considering the filtering process during policy evaluation. PILCO conduct…
Estimates system parameters from a single observation using kernel-based score.
Study on elastic curves with variable stiffness, derived from bending energy.
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…
CW-EDMD improves prediction accuracy by learning local Koopman models for different state-space regions.
Reinforcement learning is a powerful paradigm for learning optimal policies from experimental data. However, to find optimal policies, most reinforcement learning algorithms explore all possible actions, which may be harmful for real-world systems. As a consequence, learning algorithms are rarely applied on safety-crit…
Paper constructs new minimal submanifolds in spheres by spinning given ones.
Discover conservation laws from trajectories using a neural network.
In this study we introduce a new technique for symbolic regression that guarantees global optimality. This is achieved by formulating a mixed integer non-linear program (MINLP) whose solution is a symbolic mathematical expression of minimum complexity that explains the observations. We demonstrate our approach by redis…
Paper proposes adaptive control for unknown systems using reinforcement learning.
We propose new symplectic networks (SympNets) for identifying Hamiltonian systems from data based on a composition of linear, activation and gradient modules. In particular, we define two classes of SympNets: the LA-SympNets composed of linear and activation modules, and the G-SympNets composed of gradient modules. Cor…
We present a self-contained proof of the Gauss-Bonnet theorem for two-dimensional surfaces embedded in using just classical vector calculus. The exposition should be accessible to advanced undergraduate and non-expert graduate students. It may be viewed as an illustration and exercise in multivariate calculus and…
We propose directed time series regression, a new approach to estimating parameters of time-series models for use in certainty equivalent model predictive control. The approach combines merits of least squares regression and empirical optimization. Through a computational study involving a stochastic version of a well …
We study the determination of the second-order normal form for perturbed Hamiltonians , relative to the periodic flow of the unperturbed Hamiltonian . The formalism presented here is global, and can be easily implemented in any CAS. We illustrate it by means of two examples: the H…
Model-based reinforcement learning methods typically learn models for high-dimensional state spaces by aiming to reconstruct and predict the original observations. However, drawing inspiration from model-free reinforcement learning, we propose learning a latent dynamics model directly from rewards. In this work, we int…
Comparison of UQ methods in deep learning for a simple physical system.
New AI learns like neurons, generalizing from sparse rewards.
Almost toric manifolds form a class of singular Lagrangian fibered symplectic manifolds that is a natural generalization of toric manifolds. Notable examples include the K3 surface, the phase space of the spherical pendulum and rational balls useful for symplectic surgeries. The main result of the paper is a complete c…
We consider magnetic geodesic flows of the normal metrics on a class of homogeneous spaces, in particular (co)adjoint orbits of compact Lie groups. We give the proof of the non-commutative integrability of flows and show, in addition, for the case of (co)adjoint orbits, the usual Liouville integrability by means of ana…
The present paper extends the classical second-order variational problem of Herglotz type to the more general context of the Euclidean sphere S^n following variational and optimal control approaches. The relation between the Hamiltonian equations and the generalized Euler-Lagrange equations is established. This problem…
Model learns and plans in real-time under constraints for robotic systems.
Study of rotation angles in a rotating disc model.
This paper detects Markov violations in RL with noise, improving policy development.
We briefly review the notion of second order constrained (continuous) system (SOCS) and then propose a discrete time counterpart of it, which we naturally call discrete second order constrained system (DSOCS). To illustrate and test numerically our model, we construct certain integrators that simulate the evolution of …
We study the motion of a particle in the hyperbolic plane (embedded in Minkowski space), under the action of a potential that depends only on one variable. This problem is the analogous to the spherical pendulum in a unidirectional force field. However, for the discussion of the hyperbolic plane one has to distinguish …
Neural SVEs model complex systems with memory, outperforming traditional methods.
Approximate 3D elastic curves with exact constraints
Policy gradient algorithms typically combine discounted future rewards with an estimated value function, to compute the direction and magnitude of parameter updates. However, for most Reinforcement Learning tasks, humans can provide additional insight to constrain the policy learning. We introduce a general method to i…
Learning to make decisions from observed data in dynamic environments remains a problem of fundamental importance in a number of fields, from artificial intelligence and robotics, to medicine and finance. This paper concerns the problem of learning control policies for unknown linear dynamical systems so as to maximize…