Paper analyzes convergence rates of two time-scale AC and NAC algorithms.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
This work analyzes actor-critic methods for faster convergence.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
Improved bounds for non-linear SA with fast convergence.
Gradient-based temporal difference (GTD) algorithms are widely used in off-policy learning scenarios. Among them, the two time-scale TD with gradient correction (TDC) algorithm has been shown to have superior performance. In contrast to previous studies that characterized the non-asymptotic convergence rate of TDC only…
Aims to describe neural network training dynamics using two-time-scale models.
New bounds for SA with arbitrary norm contractions and Markovian noise.
In this paper, we study the problems of principal Generalized Eigenvector computation and Canonical Correlation Analysis in the stochastic setting. We propose a simple and efficient algorithm, Gen-Oja, for these problems. We prove the global convergence of our algorithm, borrowing ideas from the theory of fast-mixing M…
Motivated by their broad applications in reinforcement learning, we study the linear two-time-scale stochastic approximation, an iterative method using two different step sizes for finding the solutions of a system of two equations. Our main focus is to characterize the finite-time complexity of this method under time-…
New analysis of stochastic approximation with non-expansive mappings.
We study two time-scale linear stochastic approximation algorithms, which can be used to model well-known reinforcement learning algorithms such as GTD, GTD2, and TDC. We present finite-time performance bounds for the case where the learning rate is fixed. The key idea in obtaining these bounds is to use a Lyapunov fun…
Sharp pseudospectral bounds prevent transient amplification in coupled gradient descent.
Paper analyzes convergence of two time-scale stochastic approximation using martingale approach.
Improved stochastic approximation method reduces residual error.
We consider nonconvex-concave minimax problems, , where is nonconvex in but concave in and is a convex and bounded set. One of the most popular algorithms for solving this problem is the celebrated…
Generative Adversarial Networks (GANs) excel at creating realistic images with complex models for which maximum likelihood is infeasible. However, the convergence of GAN training has still not been proved. We propose a two time-scale update rule (TTUR) for training GANs with stochastic gradient descent on arbitrary GAN…
Circadian rhythms influence multiple essential biological activities including sleep, performance, and mood. The dim light melatonin onset (DLMO) is the gold standard for measuring human circadian phase (i.e., timing). The collection of DLMO is expensive and time-consuming since multiple saliva or blood samples are req…
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise. In particular, both the faster and slower recursions have non-additive controlled Markov noise components in addition to martingale difference noise. We analyze the asymptotic…
The paper analyzes Indian stock sectors using multifractal analysis for long and short-term investment.
During this last decades, several attempts to construct slow invariant manifold of the Lorenz-Krishnamurthy five-mode model of slow-fast interactions in the atmosphere have been made by various authors. Unfortunately, as in the case of many two-time scales singularly perturbed dynamical systems the various asymptotic p…
In this article, we discuss various implementation of L1 filtering in order to detect some properties of noisy signals. This filter consists of using a L1 penalty condition in order to obtain the filtered signal composed by a set of straight trends or steps. This penalty condition, which determines the number of breaks…
FedGAN trains GANs across distributed data sources with reduced communication.
We revisit the index leverage effect, that can be decomposed into a volatility effect and a correlation effect. We investigate the latter using a matrix regression analysis, that we call `Principal Regression Analysis' (PRA) and for which we provide some analytical (using Random Matrix Theory) and numerical benchmarks.…
Urban transformations within large and growing metropolitan areas often generate critical dynamics affecting social interactions, transport connectivity and income flow distribution. We develop a statistical-mechanical model of urban transformations, exemplified for Greater Sydney, and derive a thermodynamic descriptio…
In our model, traders interact with each other and with a central bank; they are taxed on the money they make, some of which is dissipated away by corruption. A generic feature of our model is that the richest trader always wins by 'consuming' all the others: another is the existence of a threshold wealth, below wh…
Develops a new risk measure for Markov chains' asymptotic behavior.
We investigate the Heston model with stochastic volatility and exponential tails as a model for the typical price fluctuations of the Brazilian São Paulo Stock Exchange Index (IBOVESPA). Raw prices are first corrected for inflation and a period spanning 15 years characterized by memoryless returns is chosen for the ana…
Dual training method for EBMs with overparametrized neural networks.
Atoms and molecules are important conceptual entities we invented to understand the physical world around us. The key to their usefulness lies in the organization of nuclear and electronic degrees of freedom into a single dynamical variable whose time evolution we can better imagine. The use of such effective variables…
Novel digital twin for complex systems improves performance.
Actor-critic algorithms converge to an ODE as data samples change dynamically.
We analyze stochastic approximation with Markov noise for reinforcement learning.