RL solves discrete LQ control with Gaussian optimal policy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved text generation with constraints using discrete auto-regressive biasing.
A new framework for controllable generation of discrete masked models.
We explore a new method for discrete-time control problems using randomization and entropy.
Optimizes control of noisy discrete systems without system matrix knowledge.
A new method reduces variance in training discrete latent variable models.
Efficient deep policy gradient method for continuous-time control problems.
Paper formulates mutual information optimal control for discrete-time systems.
In this paper we will discuss some new developments in the design of numerical methods for optimal control problems of Lagrangian systems on Lie groups. We will construct these geometric integrators using discrete variational calculus on Lie groups, deriving a discrete version of the second-order Euler-Lagrange equatio…
Many real-world control problems involve both discrete decision variables - such as the choice of control modes, gear switching or digital outputs - as well as continuous decision variables - such as velocity setpoints, control gains or analogue outputs. However, when defining the corresponding optimal control or reinf…
Continuous-time MBRL framework tackles control systems with Bayesian ODEs.
SOM-VQ tokenizes discrete models with semantic structure and navigable topology.
The least squares Monte Carlo algorithm has become popular for solving portfolio optimization problems. A simple approach is to approximate the value functions on a discrete grid of portfolio weights, then use control regression to generalize the discrete estimates. However, the classical global control regression can …
Improves gradient estimation for discrete distributions with variance reduction techniques.
Unified framework extends adjoint Schrödinger bridge sampler to discrete spaces.
Study scaling limits of utility indifference prices in discretized Bachelier model.
We consider a class of discrete time stochastic control problems motivated by some financial applications. We use a pathwise stochastic control approach to provide a dual formulation of the problem. This enables us to develop a numerical technique for obtaining an estimate of the value function which improves on purely…
Paper develops PAC-Bayes bounds for unknown linear systems.
A simpler edge-based discretization method without dual volumes.
The paper proposes a new stochastic intervention control model conducted in various commodity and stock markets. The essence of the phenomenon of intervention is described in accordance with current economic theory. A review of papers on intervention research has been made. A general construction of the stochastic inte…
This paper studies the properties of discrete time stochastic optimal control problems associated with portfolio selection. We investigate if optimal continuous time strategies can be used effectively for a discrete time market after a straightforward discretization. We found that Merton's strategy approximates the per…
Optimal control in latent factor models uses Tsallis entropy for exploration.
New insights into RL efficiency from managing time discretization.
In this introductory paper, we discuss how quantitative finance problems under some common risk factor dynamics for some common instruments and approaches can be formulated as time-continuous or time-discrete forward-backward stochastic differential equations (FBSDE) final-value or control problems, how these final val…
New boundary and point constraints for controlling conformal surfaces.
The paper tackles exact linearization and control of flat discrete-time systems.
Hybrid RL method optimizes trading by balancing continuous and discrete actions.
Reinforcement learning (RL) in discrete action space is ubiquitous in real-world applications, but its complexity grows exponentially with the action-space dimension, making it challenging to apply existing on-policy gradient based deep RL algorithms efficiently. To effectively operate in multidimensional discrete acti…
New estimators reduce variance in training variational autoencoders with discrete latent variables.
Safety filter for unknown discrete-time systems with learned models and noise covariance.
To better understand and improve the behavior of neural networks, a recent line of works bridged the connection between ordinary differential equations (ODEs) and deep neural networks (DNNs). The connections are made in two folds: (1) View DNN as ODE discretization; (2) View the training of DNN as solving an optimal co…
We present a framework for learning disentangled and interpretable jointly continuous and discrete representations in an unsupervised manner. By augmenting the continuous latent distribution of variational autoencoders with a relaxed discrete distribution and controlling the amount of information encoded in each latent…
For controlled discrete-time stochastic processes we introduce a new class of dynamic risk measures, which we call process-based. Their main features are that they measure risk of processes that are functions of the history of a base process. We introduce a new concept of conditional stochastic time consistency and we …
New method controls gradient error for sparse MRFs.
Paper proposes adaptive control for unknown systems using reinforcement learning.
Trading strategy mimics optimal control with simple heuristic.
The paper proves that linearization along trajectories preserves flatness in discrete-time systems.
Stochastic control problems in finance often involve complex controls at discrete times. As a result numerically solving such problems, for example using methods based on partial differential or integro-differential equations, inevitably give rise to low order accuracy, usually at most second order. In many cases one c…
Replica exchange Langevin diffusion accelerates nonconvex optimization.
Let G be a group and let M be a CAT(0) proper metric space (e.g. a simply connected complete Riemannian manifold of non-positive sectional curvature or a locally finite tree). Isometric actions of G on M are (by definition) points in the space R := Hom(G, Isom(M)) with the compact open topology. Sample theorems: 1. The…
In this paper, we consider a generalization of variational calculus which allows us to consider in the same framework different cases of mechanical systems, for instance, Lagrangian mechanics, Hamiltonian mechanics, systems subjected to constraints, optimal control theory and so on. This generalized variational calculu…
MDNS generates samples from complex discrete distributions efficiently.
Hybrid systems are characterized by having an interaction between continuous dynamics and discrete events. The contribution of this paper is to provide hybrid systems with a novel geometric formulation so that controls can be added. Using this framework we describe some new global controllability tests for hybrid contr…
Learning in models with discrete latent variables is challenging due to high variance gradient estimators. Generally, approaches have relied on control variates to reduce the variance of the REINFORCE estimator. Recent work (Jang et al. 2016, Maddison et al. 2016) has taken a different approach, introducing a continuou…
RL applied to TCLs for power consumption control.
Paper introduces a new method for risk-sensitive investment management using RL.
In this paper we propose a new methodology for solving an uncertain stochastic Markovian control problem in discrete time. We call the proposed methodology the adaptive robust control. We demonstrate that the uncertain control problem under consideration can be solved in terms of associated adaptive robust Bellman equa…
In this paper we study a new reinforcement learning setting where the environment is non-rewarding, contains several possibly related objects of various controllability, and where an apt agent Bob acts independently, with non-observable intentions. We argue that this setting defines a realistic scenario and we present …