Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4148291,2431,657 · Jun 202019922001200920172026
48 results for vanishing learning rate

Study on L2-boosting behavior as learning rate approaches zero.

problem Understanding the asymptotic behavior of L2-boosting algorithms with vanishing learning rates.
method Analyzes L2-boosting for regression with linear base learners, proving a deterministic limit and characterizing it as a solution to a linear differential equation.
result Proves the existence of a unique solution to the limit problem and analyzes the training and test error.

Study the properties of SGD in non-vanishing learning rate regime.

problem Understanding the noise and fluctuation in SGD with finite learning rates.
method Derive exact solvable results for discrete-time SGD in quadratic loss functions.
result Fluctuation caused by discrete-time dynamics is larger than continuous-time theory predicts.

Gradient amplification boosts deep learning model performance without increasing training time.

problem Vanishing gradients in deep neural networks.
method Gradient amplification approach to prevent vanishing gradients and training strategy to enable/disable across epochs.
result Improves performance of deep learning models with reduced training time.

Improved sampling for Bayesian neural networks reduces vanishing acceptance rates and increases predictive accuracy.

problem Sampling inefficiency in Bayesian neural networks, especially with deep architectures and large datasets.
method Approximate blocked Gibbs sampling to partition and sample subgroups of parameters.
result Increased predictive accuracy and quantification of predictive uncertainty in classification tasks.

Infinitesimal gradient boosting is a new algorithm derived from gradient boosting.

problem Improving the efficiency and smoothness of gradient boosting.
method Introduced a new class of randomized regression trees and used a limit process in vanishing-learning-rate asymptotic.
result Convergence of the stochastic algorithm and characterization of the limiting procedure as a unique solution of a nonlinear ODE.

The Weil-Petersson metric for the moduli space of Riemann surfaces has negative sectional curvature. Surfaces represented in the complement of a compact set in the moduli space have short geodesics. At such surfaces the Weil-Petersson metric is approximately a product metric. An almost product metric has sections with …

2019-08-26abs ↗pdf ↗

Improved convergence rates for MLE in mixture models using penalized log-likelihood.

problem Convergence rates for MLE in finite mixture models.
method Penalizing log-likelihood to discourage vanishing mixing weights, using Wasserstein distance and new loss functions.
result Improved convergence rates for some mixture components, faster than traditional methods.

New findings cast doubt on the role of λmaxλ_{max} in generalizing neural networks.

problem The role of λmaxλ_{max} in neural network generalization remains unclear.
method Experiments with various training interventions and batch sizes.
result Generalization benefits can vanish at larger batch sizes, challenging the role of λmaxλ_{max}.

Ghost mechanism explains abrupt learning in RNNs, revealing constraints on optimization landscapes.

problem Understanding abrupt learning in recurrent neural networks (RNNs) trained on working memory tasks.
method Introducing the ghost mechanism, a process driven by saddle-node bifurcations, to analyze and model abrupt learning.
result A critical learning rate scales as an inverse power law with the timescale of computation, leading to vanishing and oscillatory gradients.

Active data collection improves convergence rates in operator learning.

problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.

The paper analyzes reg-SGD for convex problems, proving convergence and quantifying the rate of convergence.

problem Minimizing convex, L-smooth functions in a Hilbert space.
method Regularized stochastic gradient descent with decaying regularization.
result Strong convergence to the minimum-norm solution without boundedness assumptions.

Paper proves multiplicative weight updates can train neural networks without learning rate tuning.

problem Vanishing and exploding gradients in gradient descent for compositional functions.
method Proves descent lemma for compositional functions using multiplicative weight updates and derives Madam optimizer.
result Madam optimizer trains state-of-the-art neural networks without learning rate tuning.

We formulate and study a general family of (continuous-time) stochastic dynamics for accelerated first-order minimization of smooth convex functions. Building on an averaging formulation of accelerated mirror descent, we propose a stochastic variant in which the gradient is contaminated by noise, and study the resultin…

2017-07-19abs ↗pdf ↗

In this paper, we formalise order-robust optimisation as an instance of online learning minimising simple regret, and propose Vroom, a zero'th order optimisation algorithm capable of achieving vanishing regret in non-stationary environments, while recovering favorable rates under stochastic reward-generating processes.…

2019-10-09abs ↗pdf ↗

Square metrics F=(α+β)2αF=\frac{(α+β)^2}α are a special class of Finsler metrics. It is the rate kind of metric category to be of excellent geometrical properties. In this paper, we discuss the so-called singular square metrics F=(bα+β)2αF=\frac{(bα+β)^2}α. A characterization for such metrics to be of vanishing Douglas curvature is p…

2016-10-31abs ↗pdf ↗

Deep linear networks minimize sharpness, avoiding large eigenvalues.

problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.

In high dimensions, the mean and geometric median are nearly identical.

problem Understanding the relationship between mean and geometric median in high-dimensional spaces.
method Analytical derivation and simulation of the distance between mean and geometric median.
result The distance between mean and geometric median vanishes with dimensionality in high dimensions.

Large learning rates work surprisingly well in standard parameterization, contrary to theory.

problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.

Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …

2018-05-23abs ↗pdf ↗

The paper explores growth rates and Perron numbers in Coxeter systems with low-dimensional Davis complexes.

problem Investigating growth rates and specific types of numbers in Coxeter systems with Davis complexes of low dimension.
method Examining Coxeter systems with Davis complexes of dimension at most 2, focusing on growth rates and specific types of numbers.
result The growth rate of Coxeter systems with Davis complexes of dimension at most 2 are either Salem or Pisot numbers, depending on the Euler characteristic.

Optimal learning rates decay to zero in easy tasks and maintain a warmup phase in hard tasks.

problem Optimizing learning rates under functional scaling laws for model training.
method Deriving optimal learning-rate schedules based on exponents ss and ββ.
result Sharp phase transition between easy and hard tasks, with different decay behaviors.

We give another proof for a result of Brick stating that the simple connectivity at infinity is a geometric property of finitely presented groups. This allows us to define the rate of vanishing of $\p1i$ for those groups which are simply connected at infinity. Further we show that this rate is linear for cocompact latt…

2002-09-02abs ↗pdf ↗

This paper investigates the efficacy of a regularized multi-task learning (MTL) framework based on SVM (M-SVM) to answer whether MTL always provides reliable results and how MTL outperforms independent learning. We first find that M-SVM is Bayes risk consistent in the limit of large sample size. This implies that despi…

2018-05-31abs ↗pdf ↗

New methods optimize training VQAs without barren plateaus, improving efficiency and applicability.

problem Barren plateaus in training variational quantum algorithms.
method Derive adaptive learning rates and use Gaussian kernels to optimize movement in parameter space.
result Optimized training methods outperform other routines and can train VQAs free of barren plateaus.

Step decay schedules improve convergence in non-convex optimization.

problem Improving convergence in non-convex optimization problems.
method Analyzing convergence rates of step decay schedules in non-convex, convex, and strongly convex problems.
result Step decay schedules achieve O(lnT/T)\mathcal{O}(\ln T/\sqrt{T}) convergence rates in various optimization scenarios.

The paper develops methods to accurately locate change points in high-dimensional mean shift models.

problem Locating change points in high-dimensional mean shift models.
method Locally refitted least squares estimator, component-wise and simultaneous rates of estimation.
result Asymptotic validity of component-wise and simultaneous confidence intervals for change point parameters.

A novel hierarchical Bayesian approach to Federated Learning reduces data exposure and improves convergence rates.

problem Data privacy and convergence in Federated Learning.
method Hierarchical Bayesian modeling and block-coordinate descent optimization.
result The proposed algorithm converges to an optimal solution with a rate of O(1/t)O(1/\sqrt{t}) and guarantees vanishing generalization error.

The paper provides theoretical guarantees for optimized sampling in compressed sensing, showing error vanishes with more measurements.

problem Theoretical and practical improvements in compressed sensing with optimized sampling schemes.
method Theoretical analysis and empirical experiments with optimized sampling schemes for subsampled unitary matrices.
result The error caused by measurement noise vanishes with an increasing number of measurements for optimized sampling schemes, assuming Gaussian noise.

We derive explicit formulas for time decay, for the European call and put options at expiry, and use them to calculate analytical approximations to the price of the American put and early exercise boundary near expiry. We show that for many families of non-Gaussian processes used in empirical studies of financial marke…

2004-04-05abs ↗pdf ↗

The paper explores how dynamic preconditioning affects the CLT in online averaging.

problem When does dynamic preconditioning preserve the Polyak-Ruppert CLT?
method The authors decompose the averaged error and identify a stabilization-rate threshold for the CLT to hold.
result The CLT holds if the dynamic remainder vanishes in L2L^2 and the stabilization rate exceeds a threshold.

Actor-critic algorithms converge to an ODE as data samples change dynamically.

problem Challenging to mathematically analyze due to non-i.i.d. data samples.
method Proved convergence to an ODE using time rescaling and geometric ergodicity.
result Convergence to the ODE limit and its properties proven.

Analyzes convergence rates for Gaussian-gated MoE model.

problem Theoretical understanding of Gaussian-gated MoE model is incomplete.
method Maximum likelihood estimation with novel Voronoi loss functions.
result MLE has distinct behaviors under different settings of Gaussian gating function parameters.

Study shows how learning and analytical models affect reneging and jockeying in a dual M/M/1 system.

problem How do learning and analytical models affect reneging and jockeying in a dual M/M/1 system?
method Analytical and online trained actor-critic models were used to study reneging and jockeying in a dual M/M/1 system.
result Both analytical and online trained actor-critic models yield the same asymptotic limits for reneging and jockeying, but differ in practical sizes.