Paper proposes a new method to stabilize noisy gradient algorithms.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper investigates a type of instability that is linked to the greedy policy improvement in approximated reinforcement learning. We show empirically that non-deterministic policy improvement can stabilize methods like LSPI by controlling the improvements' stochasticity. Additionally we show that a suitable represe…
Designing deterministic denominators for SGLD stabilizes large drifts.
Paper explores stability, regularization, and gradient flows for stochastic inverse problems.
Unified framework for solving fixed-point equations in deterministic and stochastic settings.
Unified treatment of RC in stochastic and deterministic settings.
This paper analyzes how periodic and soft target updates stabilize linear Q-learning.
Develops a deterministic method to approximate NSDEs for better uncertainty quantification.
Bayesian explanations are more resilient to adversarial attacks than deterministic ones.
Develops DPG methods for continuous-time RL with deterministic policies.
Systemic risk measures are crucial for the stability of financial markets, yet classical formulations fail to capture the complexity of market volatility. We propose a new framework for systemic risk measurement on the variable-exponent Bochner-Lebesgue space , where the exponent is a random va…
SGD works well with large learning rates at the edge of stability.
This paper investigates trajectory tracking problem for a class of underactuated autonomous underwater vehicles (AUVs) with unknown dynamics and constrained inputs. Different from existing policy gradient methods which employ single actor-critic but cannot realize satisfactory tracking control accuracy and stable learn…
We propose a heterogeneous agent market model (HAM) in continuous time. The market is populated by fundamental traders and chartists, who both use simple linear trading rules. Most of the related literature explores stability, price dynamics and profitability either within deterministic models or by simulation. Our nov…
In deterministic optimization, line searches are a standard tool ensuring stability and efficiency. Where only stochastic gradients are available, no direct equivalent has so far been formulated, because uncertain gradients do not allow for a strict sequence of decisions collapsing the search space. We construct a prob…
In deterministic optimization, line searches are a standard tool ensuring stability and efficiency. Where only stochastic gradients are available, no direct equivalent has so far been formulated, because uncertain gradients do not allow for a strict sequence of decisions collapsing the search space. We construct a prob…
Develops a new reinforcement learning framework for complex control problems.
Local Interpretable Model-Agnostic Explanations (LIME) is a popular technique used to increase the interpretability and explainability of black box Machine Learning (ML) algorithms. LIME typically generates an explanation for a single prediction by any ML model by learning a simpler interpretable model (e.g. linear cla…
Proposes a new policy gradient algorithm to improve reinforcement learning efficiency and stability.
Unified kernel framework extends to stochastic systems, improving numerical stability.
Clustering is a crucial component of many data mining systems involving the analysis and exploration of various data. Data diversity calls for clustering algorithms to be accurate while providing stable (i.e., deterministic and robust) results on arbitrary input networks. Moreover, modern systems often operate with lar…
Framework simulates market microstructure with stable Hawkes processes.
Develops pathwise analysis for log-optimal portfolios using rough paths theory.
ZDPG learns model-free policies without critics, improving on PG.
Model predicts alternating market dominance for two competing firms.
Continuous control imitation learning fails if expert actions are smooth.
New approach to concentration inequalities for unbounded state space dynamical systems.
This study analyzes convergence and stability of reinforcement learning algorithms.
We provide existence, uniqueness and stability results for affine stochastic Volterra equations with -kernels and jumps. Such equations arise as scaling limits of branching processes in population genetics and self-exciting Hawkes processes in mathematical finance. The strategy we adopt for the existence part is b…
Econometrics is based on the nonempiric notion of utility. Prices, dynamics, and market equilibria are supposed to be derived from utility. Utility is usually treated by economists as a price potential, other times utility rates are treated as Lagrangians. Assumptions of integrability of Lagrangians and dynamics are im…
The paper analyzes the stationarity of stochastic Volterra integral equations and introduces fake stationary regimes.
In this paper, we present our approach to solve a physics-based reinforcement learning challenge "Learning to Run" with objective to train physiologically-based human model to navigate a complex obstacle course as quickly as possible. The environment is computationally expensive, has a high-dimensional continuous actio…
Optimism stabilizes Thompson Sampling for adaptive inference in multi-armed bandits.
We introduce a new recursive aggregation procedure called Bernstein Online Aggregation (BOA). The exponential weights include an accuracy term and a second order term that is a proxy of the quadratic variation as in Hazan and Kale (2010). This second term stabilizes the procedure that is optimal in different senses. We…
We consider the matrix completion problem with a deterministic pattern of observed entries. In this setting, we aim to answer the question: under what condition there will be (at least locally) unique solution to the matrix completion problem, i.e., the underlying true matrix is identifiable. We answer the question fro…
We show a concise extension of the monotone stability approach to backward stochastic differential equations (BSDEs) that are jointly driven by a Brownian motion and a random measure for jumps, which could be of infinite activity with a non-deterministic and time inhomogeneous compensator. The BSDE generator function c…
Bayesian MoE framework improves LLMs' uncertainty detection.
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
Online algorithms stabilize in feedback loops of performative prediction.
Chatter identification and detection in machining processes has been an active area of research in the past two decades. Part of the challenge in studying chatter is that machining equations that describe its occurrence are often nonlinear delay differential equations. The majority of the available tools for chatter id…
Paper uses SAC and DDPG to optimize cryptocurrency portfolios.
Randomness is crucial for stability in learning and statistics, especially for differential privacy.
Deep learning has become an area of interest in most scientific areas, including physical sciences. Modern networks apply real-valued transformations on the data. Particularly, convolutions in convolutional neural networks discard phase information entirely. Many deterministic signals, such as seismic data or electrica…
Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is analyzed in solving the problem. The convergence of the iterations and the optim…
The need for parameter estimation with massive datasets has reinvigorated interest in stochastic optimization and iterative estimation procedures. Stochastic approximations are at the forefront of this recent development as they yield procedures that are simple, general, and fast. However, standard stochastic approxima…
This paper studies Thompson sampling's arm-pull dynamics and inference, revealing key differences from UCB algorithms.
Extracts factors from Treasury yields using ML techniques.
In the present paper and the companion paper [8] a probabilistic (statistical mechanical) approach to the study of canonical metrics and measures on a complex algebraic variety X is introduced. On any such variety with positive Kodaira dimension a canonical (birationally invariant) random point processes is defined and…