Simplifies RL training with fewer techniques, reducing bias and instability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Venice used 'helicopter money' to subsidize during famine and plague, but it caused instability.
Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy () and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The var…
We extend in a minimal way the stylized model introduced in in "Tipping Points in Macroeconomic Agent Based Models" [JEDC 50, 29-61 (2015)], with the aim of investigating the role and efficacy of monetary policy of a `Central Bank' that sets the interest rate such as to steer the economy towards a prescribed inflation …
This paper investigates a type of instability that is linked to the greedy policy improvement in approximated reinforcement learning. We show empirically that non-deterministic policy improvement can stabilize methods like LSPI by controlling the improvements' stochasticity. Additionally we show that a suitable represe…
Following the financial crisis of 2007-2008, a deep analogy between the origins of instability in financial systems and complex ecosystems has been pointed out: in both cases, topological features of network structures influence how easily distress can spread within the system. However, in financial network models, the…
Detects adversarial directions to make reinforcement learning policies more robust.
Stabilizes policy optimization with off-policy data using divergence augmentation.
We present Multitask Soft Option Learning(MSOL), a hierarchical multitask framework based on Planning as Inference. MSOL extends the concept of options, using separate variational posteriors for each task, regularized by a shared prior. This ''soft'' version of options avoids several instabilities during training in a …
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new policy after every policy gradient update. Despite enormous success of off-policy …
Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot scale to tasks with very high state and action dimensionality such as 3D humanoid l…
We study the finite-size effects in some scaling systems, and show that the finite number of agents N leads to a cut-off in the upper value of the Pareto law for the relative individual wealth. The exponent of the Pareto law obtained in stochastic multiplicative market models is crucially affected by the fact that …
Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, particularly for on-policy methods. Previous exploration methods either rely on complex structure to estimate the novelty of states, or incur sens…
New RL algorithm ensures stable, replicable policies.
PPOS improves PPO by smoothing the surrogate objective function.
Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution, and can make only …
Adapts model-based advice to stabilize black-box policies for nonlinear control.
In many real-world reinforcement learning applications, access to the environment is limited to a fixed dataset, instead of direct (online) interaction with the environment. When using this data for either evaluation or training of a new policy, accurate estimates of discounted stationary distribution ratios -- correct…
Financial system being the place of metting capital flows (equality between saving and investment), a volatility of capital flows can destroy the robustness and good working of financial system, it means subvert financial stability. The same a weak financial system, few regulated and bad manage can exacerbate volatilit…
DisCor corrects reinforcement learning issues by re-weighting collected data.
The paper explores how AI trading agents' similar information representation can cause financial market instability.
New method improves causal effect estimation by addressing imbalance in training data.
With negative growth in real production in many countries and debt levels which become an increasing burden on developed societies, the calls for a change in economic policy and even the monetary system become louder and increasingly impatient. We research the consequences of a system of credit and debt, that still all…
BCPO optimizes offline RL policies by converting uncertainty into conservative bounds.
New algorithms improve reinforcement learning stability and performance.
A new reinforcement learning method uses model derivatives to improve policy optimization.
Proximal policy optimization (PPO) is one of the most successful deep reinforcement-learning methods, achieving state-of-the-art performance across a wide range of challenging tasks. However, its optimization behavior is still far from being fully understood. In this paper, we show that PPO could neither strictly restr…
SUNRISE improves off-policy RL algorithms by integrating ensemble methods.
Many complex domains, such as robotics control and real-time strategy (RTS) games, require an agent to learn a continuous control. In the former, an agent learns a policy over and in the latter, over a discrete set of actions each of which is parametrized by a continuous parameter. Such problems are natu…
In recent years, there has been significant progress in applying deep reinforcement learning (RL) for solving challenging problems across a wide variety of domains. Nevertheless, convergence of various methods has been shown to suffer from inconsistencies, due to algorithmic instability and variance, as well as stochas…
Investment disputes increase stock volatility, especially for companies with negative outcomes.
A new risk measure (FRM) for EM FI returns helps investors protect against volatility and policy instability.
Previously, the exploding gradient problem has been explained to be central in deep learning and model-based reinforcement learning, because it causes numerical issues and instability in optimization. Our experiments in model-based reinforcement learning imply that the problem is not just a numerical issue, but it may …
Gradient-based meta-RL fails with incorrect task distributions, leading to instability and poor performance.
We improve current instability-based methods for the selection of the number of clusters in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously unaccounted driver of cluster instability. We show that our normalized instability mea…
ADPO optimizes relative advantage in reinforcement learning from human feedback.
New method improves deep policy gradient algorithms by learning relative state values.
An optimal feedback controller for a given Markov decision process (MDP) can in principle be synthesized by value or policy iteration. However, if the system dynamics and the reward function are unknown, a learning agent must discover an optimal controller via direct interaction with the environment. Such interactive d…
Modeling HFT interactions reveals market instability.
New method improves stability of soft FQI for offline RL.
MODULE solves LfO problem with high sample efficiency and stability.
New proof of instability for certain Einstein metrics.
New RL method learns from state transitions without actions.
Study stability and instability of Ricci-flat metrics under generalized Ricci flow.
Some exotic compact objects possess evanescent ergosurfaces: timelike submanifolds on which a Killing vector field, which is timelike everywhere else, becomes null. We show that any manifold possessing an evanescent ergosurface but no event horizon exhibits a linear instability of a peculiar kind: either there are solu…
Off-policy learning is powerful for reinforcement learning. However, the high variance of off-policy evaluation is a critical challenge, which causes off-policy learning falls into an uncontrolled instability. In this paper, for reducing the variance, we introduce control variate technique to $\math…
Training intelligent agents through reinforcement learning is a notoriously unstable procedure. Massive parallelization on GPUs and distributed systems has been exploited to generate a large amount of training experiences and consequently reduce instabilities, but the success of training remains strongly influenced by …
Interval Neural Networks detect instabilities in image reconstructions.