Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

74148221295 · Jun 202019922001200920172026
48 results for vanishing rate

Improved sampling for Bayesian neural networks reduces vanishing acceptance rates and increases predictive accuracy.

problem Sampling inefficiency in Bayesian neural networks, especially with deep architectures and large datasets.
method Approximate blocked Gibbs sampling to partition and sample subgroups of parameters.
result Increased predictive accuracy and quantification of predictive uncertainty in classification tasks.

The Weil-Petersson metric's curvature vanishes for surfaces with short geodesics.

problem Understanding the vanishing rate of Weil-Petersson sectional curvatures.
method Analyzing the moduli space of Riemann surfaces, bounding curvature, and using product metrics.
result Bounding the sectional curvature away from zero in terms of geodesic lengths.

Study on L2-boosting behavior as learning rate approaches zero.

problem Understanding the asymptotic behavior of L2-boosting algorithms with vanishing learning rates.
method Analyzes L2-boosting for regression with linear base learners, proving a deterministic limit and characterizing it as a solution to a linear differential equation.
result Proves the existence of a unique solution to the limit problem and analyzes the training and test error.

Study the properties of SGD in non-vanishing learning rate regime.

problem Understanding the noise and fluctuation in SGD with finite learning rates.
method Derive exact solvable results for discrete-time SGD in quadratic loss functions.
result Fluctuation caused by discrete-time dynamics is larger than continuous-time theory predicts.

Improved convergence rates for MLE in mixture models using penalized log-likelihood.

problem Convergence rates for MLE in finite mixture models.
method Penalizing log-likelihood to discourage vanishing mixing weights, using Wasserstein distance and new loss functions.
result Improved convergence rates for some mixture components, faster than traditional methods.

Paper analyzes if GANs with adversarial features outperform standard supervised learning.

problem Whether adversarial features improve performance over standard supervised learning.
method Theoretical analysis and empirical risk comparison.
result Supervised learning with adversarial features can outperform sole supervision under certain conditions.

Study on solutions of degenerate equations on manifolds, linking behavior to geometry and decay rates.

problem Behavior of solutions to degenerate parabolic equations on manifolds with inhomogeneous density.
method Analysis of Cauchy problem on Riemannian manifolds, considering weight function as capacitary coefficient.
result Estimates of vanishing rate and finite speed of propagation in subcritical ranges, universal bounds and blow-up in supercritical ranges.

TSGO optimizes gradients in tensor networks to avoid vanishing/exploding issues.

problem Gradient vanishing and exploding problems in deep learning models.
method TSGO rotates parameters towards gradient direction in tangent space of normalized state.
result TSGO naturally determines learning rate based on angle between parameters and gradient.

Gradient amplification boosts deep learning model performance without increasing training time.

problem Vanishing gradients in deep neural networks.
method Gradient amplification approach to prevent vanishing gradients and training strategy to enable/disable across epochs.
result Improves performance of deep learning models with reduced training time.

Square metrics F=(α+β)2αF=\frac{(α+β)^2}α are a special class of Finsler metrics. It is the rate kind of metric category to be of excellent geometrical properties. In this paper, we discuss the so-called singular square metrics F=(bα+β)2αF=\frac{(bα+β)^2}α. A characterization for such metrics to be of vanishing Douglas curvature is p…

2016-10-31abs ↗pdf ↗

Infinitesimal gradient boosting is a new algorithm derived from gradient boosting.

problem Improving the efficiency and smoothness of gradient boosting.
method Introduced a new class of randomized regression trees and used a limit process in vanishing-learning-rate asymptotic.
result Convergence of the stochastic algorithm and characterization of the limiting procedure as a unique solution of a nonlinear ODE.

The paper explores growth rates and Perron numbers in Coxeter systems with low-dimensional Davis complexes.

problem Investigating growth rates and specific types of numbers in Coxeter systems with Davis complexes of low dimension.
method Examining Coxeter systems with Davis complexes of dimension at most 2, focusing on growth rates and specific types of numbers.
result The growth rate of Coxeter systems with Davis complexes of dimension at most 2 are either Salem or Pisot numbers, depending on the Euler characteristic.

The paper analyzes reg-SGD for convex problems, proving convergence and quantifying the rate of convergence.

problem Minimizing convex, L-smooth functions in a Hilbert space.
method Regularized stochastic gradient descent with decaying regularization.
result Strong convergence to the minimum-norm solution without boundedness assumptions.

New method finds optimal learning rates for neural nets.

problem Finding optimal learning rates in stochastic neural networks.
method Gradient-only line searches using Non-negative Associative Gradient Projection Points (NN-GPPs).
result Learning rates can be reliably resolved as step sizes along search directions.

We give another proof for a result of Brick stating that the simple connectivity at infinity is a geometric property of finitely presented groups. This allows us to define the rate of vanishing of $\p1i$ for those groups which are simply connected at infinity. Further we show that this rate is linear for cocompact latt…

2002-09-02abs ↗pdf ↗

New findings cast doubt on the role of λmaxλ_{max} in generalizing neural networks.

problem The role of λmaxλ_{max} in neural network generalization remains unclear.
method Experiments with various training interventions and batch sizes.
result Generalization benefits can vanish at larger batch sizes, challenging the role of λmaxλ_{max}.

We formulate and study a general family of (continuous-time) stochastic dynamics for accelerated first-order minimization of smooth convex functions. Building on an averaging formulation of accelerated mirror descent, we propose a stochastic variant in which the gradient is contaminated by noise, and study the resultin…

2017-07-19abs ↗pdf ↗

The paper develops methods to accurately locate change points in high-dimensional mean shift models.

problem Locating change points in high-dimensional mean shift models.
method Locally refitted least squares estimator, component-wise and simultaneous rates of estimation.
result Asymptotic validity of component-wise and simultaneous confidence intervals for change point parameters.

The paper provides theoretical guarantees for optimized sampling in compressed sensing, showing error vanishes with more measurements.

problem Theoretical and practical improvements in compressed sensing with optimized sampling schemes.
method Theoretical analysis and empirical experiments with optimized sampling schemes for subsampled unitary matrices.
result The error caused by measurement noise vanishes with an increasing number of measurements for optimized sampling schemes, assuming Gaussian noise.

The study bounds growth of Hodge numbers and computes L2L^2-Betti numbers for irregular varieties.

problem Bounding growth of normalized Hodge numbers and computing L2L^2-Betti numbers for irregular varieties.
method Analysis of abelian covers, weak generic Nakano vanishing theorem, and convergence of plurigenera.
result Optimal bounds on the growth of normalized Hodge numbers and computation of L2L^2-Betti numbers.

Active data collection improves convergence rates in operator learning.

problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.

We derive explicit formulas for time decay, for the European call and put options at expiry, and use them to calculate analytical approximations to the price of the American put and early exercise boundary near expiry. We show that for many families of non-Gaussian processes used in empirical studies of financial marke…

2004-04-05abs ↗pdf ↗

The paper explores how dynamic preconditioning affects the CLT in online averaging.

problem When does dynamic preconditioning preserve the Polyak-Ruppert CLT?
method The authors decompose the averaged error and identify a stabilization-rate threshold for the CLT to hold.
result The CLT holds if the dynamic remainder vanishes in L2L^2 and the stabilization rate exceeds a threshold.

Ghost mechanism explains abrupt learning in RNNs, revealing constraints on optimization landscapes.

problem Understanding abrupt learning in recurrent neural networks (RNNs) trained on working memory tasks.
method Introducing the ghost mechanism, a process driven by saddle-node bifurcations, to analyze and model abrupt learning.
result A critical learning rate scales as an inverse power law with the timescale of computation, leading to vanishing and oscillatory gradients.

Paper proves multiplicative weight updates can train neural networks without learning rate tuning.

problem Vanishing and exploding gradients in gradient descent for compositional functions.
method Proves descent lemma for compositional functions using multiplicative weight updates and derives Madam optimizer.
result Madam optimizer trains state-of-the-art neural networks without learning rate tuning.

We extend a recent synchronization analysis of exact finite-state sources to nonexact sources for which synchronization occurs only asymptotically. Although the proof methods are quite different, the primary results remain the same. We find that an observer's average uncertainty in the source state vanishes exponential…

2010-11-06abs ↗pdf ↗

In high dimensions, the mean and geometric median are nearly identical.

problem Understanding the relationship between mean and geometric median in high-dimensional spaces.
method Analytical derivation and simulation of the distance between mean and geometric median.
result The distance between mean and geometric median vanishes with dimensionality in high dimensions.

The paper constructs singularities for Lagrangian flow in Gibbons-Hawking spaces with vanishing mean curvature.

problem Infinite-time singularities with vanishing mean curvature for Lagrangian mean curvature flow in Gibbons-Hawking spaces.
method One-parameter family of barrier curves and detailed asymptotic analysis.
result The mean curvature converges uniformly to zero, but the second fundamental form becomes unbounded.

Calculates winning probability for three candidates based on support rates and information timing.

problem Determining optimal strategy for three candidates in an election.
method Closed-form solution using support rates, political spectrum positioning, time left, and information revelation rate.
result Optimal strategy can be complex, especially for candidates in the center of a polarized electorate.

Vanishing long-term gradients are a major issue in training standard recurrent neural networks (RNNs), which can be alleviated by long short-term memory (LSTM) models with memory cells. However, the extra parameters associated with the memory cells mean an LSTM layer has four times as many parameters as an RNN with the…

2018-02-22abs ↗pdf ↗

In this paper, asymptotic results in a long-term growth rate portfolio optimization model under both fixed and proportional transaction costs are obtained. More precisely, the convergence of the model when the fixed costs tend to zero is investigated. A suitable limit model with purely proportional costs is introduced …

2016-11-04abs ↗pdf ↗

Estimates change point in high dimensional time series models.

problem Change point estimation in high dimensional time series.
method Plug-in least squares estimator with sufficient conditions for adaptivity.
result Optimal rate of convergence Op(ξ2)O_p(ξ^{-2}) in integer scale.

The study examines a financial model with sticky prices and finds no arbitrage when interest rate is zero.

problem Analyzing financial markets with sticky asset prices and proving no arbitrage conditions.
method Introduced a financial market model with a risky asset following a sticky geometric Brownian motion and a riskless asset with a constant interest rate. Proved no arbitrage conditions and derived pricing equations.
result No arbitrage conditions are met only when the interest rate is zero, and all replicable payoffs are derived under this condition.

The paper establishes convergence rates for MoE models in classification problems.

problem Understanding the behavior of MoE models in classification settings.
method Established convergence rates for density and parameter estimation in softmax gating multinomial logistic MoE models.
result Parameter estimation rates are significantly improved with a novel modified softmax gating function.

AdaGrad fails to adapt to Hölder-smoothness in composite optimization problems.

problem AdaGrad's convergence rate is suboptimal for composite objectives.
method Exhibited a simple one-dimensional convex problem to highlight AdaGrad's limitations.
result AdaGrad does not achieve the classical convergence rate for Hölder-smooth objectives.

Step decay schedules improve convergence in non-convex optimization.

problem Improving convergence in non-convex optimization problems.
method Analyzing convergence rates of step decay schedules in non-convex, convex, and strongly convex problems.
result Step decay schedules achieve O(lnT/T)\mathcal{O}(\ln T/\sqrt{T}) convergence rates in various optimization scenarios.

Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …

2018-05-23abs ↗pdf ↗

We consider the problem of decentralized consensus optimization, where the sum of nn smooth and strongly convex functions are minimized over nn distributed agents that form a connected network. In particular, we consider the case that the communicated local decision variables among nodes are quantized in order to all…

2018-06-29abs ↗pdf ↗

Deep linear networks minimize sharpness, avoiding large eigenvalues.

problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.

Improved SEG method converges to Nash equilibrium in bilinear games.

problem Stochastic bilinear minimax optimization problem
method Stochastic ExtraGradient (SEG) method with constant step size, iteration averaging, and scheduled restarting.
result Provable convergence to Nash equilibrium under standard settings, optimal convergence rate in interpolation setting.