Study on fluctuations in neural network kernels and predictions, focusing on finite width effects.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Stochastic optimization algorithms with variance reduction have proven successful for minimizing large finite sums of functions. Unfortunately, these techniques are unable to deal with stochastic perturbations of input data, induced for example by data augmentation. In such cases, the objective is no longer a finite su…
Unified method for MMD variance estimation improves accuracy and computational efficiency.
New unbiased gradient estimators for complex optimization problems.
SVRN accelerates Newton methods by reducing variance and improving performance.
The paper extends confidence sequences for infinite variance data.
SignSVRG improves SignSGD by reducing variance, achieving similar convergence rates.
Study bounds variance modulation function for K-spider distributions.
Gaussian processes (GPs) offer a flexible class of priors for nonparametric Bayesian regression, but popular GP posterior inference methods are typically prohibitively slow or lack desirable finite-data guarantees on quality. We develop an approach to scalable approximate GP regression with finite-data guarantees on th…
Monotonicity of normalized implied-volatility coordinates under no-arbitrage
In this paper we propose a novel variance reduction approach for additive functionals of Markov chains based on minimization of an estimate for the asymptotic variance of these functionals over suitable classes of control variates. A distinctive feature of the proposed approach is its ability to significantly reduce th…
Paper develops momentum schemes with variance reduction for non-convex composition optimization.
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of con…
Study variance-optimal hedging of forward curve derivatives under stochastic volatility.
Paper develops methods for statistical inference in SGD with infinite variance.
Improved variance reduction for Riemannian non-convex optimization with adaptive batch size.
Paper shows MoM is optimal under adversarial contamination for certain distributions.
Simultaneous orthogonal matching pursuit (SOMP) and block OMP (BOMP) are two widely used techniques for sparse support recovery in multiple measurement vector (MMV) and block sparse (BS) models respectively. For optimal performance, both SOMP and BOMP require \textit{a priori} knowledge of signal sparsity or noise vari…
RMDA trains structured neural networks with regularization and variance reduction.
Increasing variance of losses improves learning with noisy labels.
Paper proposes a self-supervised method to denoise autoregressive signals with heavy-tailed noise.
We consider the problem of minimizing the composition of a smooth (nonconvex) function and a smooth vector mapping, where the inner mapping is in the form of an expectation over some random variable or a finite sum. We propose a stochastic composite gradient method that employs an incremental variance-reduced estimator…
We study confidence intervals based on hard-thresholding, soft-thresholding, and adaptive soft-thresholding in a linear regression model where the number of regressors may depend on and diverge with sample size . In addition to the case of known error variance, we define and study versions of the estimators when…
The paper analyzes the variance of different shuffling methods in stochastic gradient descent.
Improved mean estimation for symmetric distributions with finite-sample guarantees.
Techniques for reducing the variance of gradient estimates used in stochastic programming algorithms for convex finite-sum problems have received a great deal of attention in recent years. By leveraging dissipativity theory from control, we provide a new perspective on two important variance-reduction algorithms: SVRG …
The study reveals a transition in neural network performance from infinite-width to variance-limited behavior as dataset size increases.
New method solves root-finding problems with faster convergence.
Unified framework for decentralized optimization combining gradient tracking and variance reduction.
Before training a neural net, a classic rule of thumb is to randomly initialize the weights so the variance of activations is preserved across layers. This is traditionally interpreted using the total variance due to randomness in both weights \emph{and} samples. Alternatively, one can interpret the rule of thumb as pr…
We study Frank-Wolfe methods for nonconvex stochastic and finite-sum optimization problems. Frank-Wolfe methods (in the convex case) have gained tremendous recent interest in machine learning and optimization communities due to their projection-free property and their ability to exploit structured constraints. However,…
Improves SVGD for high-dimensional Bayesian inference by reducing variance collapse.
We propose two algorithms that can find local minima faster than the state-of-the-art algorithms in both finite-sum and general stochastic nonconvex optimization. At the core of the proposed algorithms is using stochastic nested variance reduction (Zhou et al., 2018a), which outperforms the s…
MUSE provides unbiased stopping estimates for optimal problems.
The posterior variance of Gaussian processes is a valuable measure of the learning error which is exploited in various applications such as safe reinforcement learning and control design. However, suitable analysis of the posterior variance which captures its behavior for finite and infinite number of training data is …
This paper presents a multinomial method for option pricing when the underlying asset follows an exponential Variance Gamma process. The continuous time Variance Gamma process is approximated by a discrete time Markov chain with the same firsts four cumulants. This approach is particularly convenient for pricing Americ…
The paper analyzes off-policy TD-learning using generalized Bellman operators and provides finite-sample bounds.
New algorithm reduces optimization complexity in adaptive mirror descent.
A large portfolio of independent returns is optimized under the variance risk measure with a ban on short positions. The no-short selling constraint acts as an asymmetric regularizer, setting some of the portfolio weights to zero and keeping the out of sample estimator for the variance bounded, avoiding the di…
Truncated Lévy flights are random walks in which the arbitrarily large steps of a Lévy flight are eliminated. Since this makes the variance finite, the central limit theorem applies, and as time increases the probability distribution of the increments becomes Gaussian. Here, truncated Lévy flights with correlated fluct…
Revisits Lee's Moment Formula, relaxing moment assumptions for implied volatility.
We show that on-policy policy gradient (PG) and its variance reduction variants can be derived by taking finite difference of function evaluations supplied by estimators from the importance sampling (IS) family for off-policy evaluation (OPE). Starting from the doubly robust (DR) estimator (Jiang & Li, 2016), we provid…
We study the conditions under which one is able to efficiently apply variance-reduction and acceleration schemes on finite sum optimization problems. First, we show that, perhaps surprisingly, the finite sum structure by itself, is not sufficient for obtaining a complexity bound of $\tilde{\cO}((n+L/μ)\ln(1/ε))$ for $L…
VRER selectively reuses past observations to reduce variance in policy optimization.
Meta-learning variance reduced via Laplace approximation for regression tasks.
Paper provides convergence guarantees for off-policy NAC with finite sample complexity.
We present novel minibatch stochastic optimization methods for empirical risk minimization problems, the methods efficiently leverage variance reduced first-order and sub-sampled higher-order information to accelerate the convergence speed. For quadratic objectives, we prove improved iteration complexity over state-of-…
We develop an approach to risk minimization and stochastic optimization that provides a convex surrogate for variance, allowing near-optimal and computationally efficient trading between approximation and estimation error. Our approach builds off of techniques for distributionally robust optimization and Owen's empiric…