Open manifolds with nonnegative Ricci curvature have virtually abelian fundamental groups if they escape from bounded balls at a small rate.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SGD's escape rate depends on log loss barrier, not linear loss barrier.
Geodesic loops escape from balls at a sublinear rate imply virtually abelian fundamental group.
The Dirichlet random walk on manifolds has a positive escape rate if the cover is non-amenable.
New result on group actions in CAT(0) spaces with vanishing escape rate.
A new method helps escape saddle points in non-convex optimization.
New methods help escape strict saddle points in nonsmooth optimization.
We consider minimizing a nonconvex, smooth function on a Riemannian manifold . We show that a perturbed version of Riemannian gradient descent algorithm converges to a second-order stationary point (and hence is able to escape saddle points on the manifold). The rate of convergence depends as o…
The paper analyzes how noise geometry influences the performance of SGD in machine learning.
Improves training GANs by escaping limit cycles.
This paper shows that a perturbed form of gradient descent converges to a second-order stationary point in a number iterations which depends only poly-logarithmically on dimension (i.e., it is almost "dimension-free"). The convergence rate of this procedure matches the well-known convergence rate of gradient descent to…
We present the first tree-based regressor whose convergence rate depends only on the intrinsic dimension of the data, namely its Assouad dimension. The regressor uses the RPtree partitioning procedure, a simple randomized variant of k-d trees.
This paper explains why Adam generalizes worse than SGD by analyzing its components.
We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a strong component along these directions. Furthermore, we show that - contrary to the case of isotropic noise - this variance is proportional to t…
SALR improves deep learning generalization by dynamically adjusting learning rates.
Classifies conformal transformations in spacetimes without observer horizons.
The paper analyzes neural network dynamics after weights escape the origin.
We shortly review the statistical properties of the escape times, or hitting times, for stock price returns by using different models which describe the stock market evolution. We compare the probability function (PF) of these escape times with that obtained from real market data. Afterwards we analyze in detail the ef…
Algorithm finds safe zones in policy Markov Decision Processes to limit trajectory escape.
This paper proposes a stochastic variant of a classic algorithm---the cubic-regularized Newton method [Nesterov and Polyak 2006]. The proposed algorithm efficiently escapes saddle points and finds approximate local minima for general smooth, nonconvex functions in only stochastic gradien…
Gradient descent can take exponentially long to escape saddle points in 2D.
New analysis of SGD with MCMC gradient estimator shows convergence rate and saddle point escape.
New algorithm helps escape saddle points in optimization problems.
The paper proves Zimmer's conjecture for non-uniform lattices by controlling mass escape and Lyapunov exponents.
Nesterov's accelerated gradient descent (AGD), an instance of the general family of "momentum methods", provably achieves faster convergence rate than gradient descent (GD) in the convex setting. However, whether these methods are superior to GD in the nonconvex setting remains open. This paper studies a simple variant…
Algorithm learns stochastic system dynamics from data.
Deep ReLU networks escape from the origin via saddle points with a low-rank bias.
Houdini finds high-dimensional saddle points under few constraints.
Study large deviations and speed of random walks in hyperbolic spaces.
HA-SME models SGD dynamics with Hessian info for better escaping behaviors.
Many nonparametric regressors were recently shown to converge at rates that depend only on the intrinsic dimension of data. These regressors thus escape the curse of dimension when high-dimensional data has low intrinsic dimension (e.g. a manifold). We show that k-NN regression is also adaptive to intrinsic dimension. …
Understanding the behavior of stochastic gradient descent (SGD) in the context of deep neural networks has raised lots of concerns recently. Along this line, we study a general form of gradient based optimization dynamics with unbiased noise, which unifies SGD and standard Langevin dynamics. Through investigating this …
Warning signs about the developing economic crisis in Greece were present in the growth rate of the Gross Domestic Product (GDP) and in the growth of the GDP well before the economic collapse. The growth rate was strongly unstable. On average, in less than 50 years, it decreased 10-folds but after reaching a low minimu…
Paper shows faster convergence to local-minimizers in over-parametrized models under interpolation-like conditions.
We solve the escape problem for the Heston random diffusion model. We obtain exact expressions for the survival probability (which ammounts to solving the complete escape problem) as well as for the mean exit time. We also average the volatility in order to work out the problem for the return alone regardless volatilit…
In most sampling algorithms, including Hamiltonian Monte Carlo, transition rates between states correspond to the probability of making a transition in a single time step, and are constrained to be less than or equal to 1. We derive a Hamiltonian Monte Carlo algorithm using a continuous time Markov jump process, and ar…
We study the mean escape time in a market model with stochastic volatility. The process followed by the volatility is the Cox Ingersoll and Ross process which is widely used to model stock price fluctuations. The market model can be considered as a generalization of the Heston model, where the geometric Brownian motion…
This paper proposes a new global optimization algorithm using deep learning.
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
Nonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoi…
Study of SGD with state-dependent noise, improving escape from local minima.
Although gradient descent (GD) almost always escapes saddle points asymptotically [Lee et al., 2016], this paper shows that even with fairly natural random initialization schemes and non-pathological functions, GD can be significantly slowed down by saddle points, taking exponential time to escape. On the other hand, g…
SGD with large learning rates can converge to local maxima.
While first-order optimization methods such as stochastic gradient descent (SGD) are popular in machine learning (ML), they come with well-known deficiencies, including relatively-slow convergence, sensitivity to the settings of hyper-parameters such as learning rate, stagnation at high training errors, and difficulty …
Study the properties of SGD in non-vanishing learning rate regime.
Learning rate decay (lrDecay) is a \emph{de facto} technique for training modern neural networks. It starts with a large learning rate and then decays it multiple times. It is empirically observed to help both optimization and generalization. Common beliefs in how lrDecay works come from the optimization analysis of (S…
The roundworm C. elegans exhibits robust escape behavior in response to rapidly rising temperature. The behavior lasts for a few seconds, shows history dependence, involves both sensory and motor systems, and is too complicated to model mechanistically using currently available knowledge. Instead we model the process p…
New metrics help predict Brownian motion on surfaces and higher dimensions.