Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work analyzes the statistical properties of adaptive gradient methods.
In this paper, we look for properties of gradient Yamabe solitons on top of warped product manifolds. Utilizing the maximum principle, we find lower bound estimates for both the potential function of the soliton and the scalar curvature of the warped product. By slightly modifying Li-Yau's technique so that we can hand…
Paper investigates conditions for independence of weak gradients on metric spaces.
We study some equivalent properties of the curvature-dimension conditions inequality on infinite, but locally finite graph. These equivalences are gradient estimate, Poincaré type inequalities and reverse Poincaré inequalities. And we also obtain one equivalent property of gradient estimate for a new notion o…
The staircase property aids deep learning by guiding hierarchical feature learning.
Gradient descent and its variants are widely used in machine learning. However, oracle access of gradient may not be available in many applications, limiting the direct use of gradient descent. This paper proposes a method of estimating gradient to perform gradient descent, that converges to a stationary point for gene…
The paper explores properties of projections and gradient methods in hyperbolic space forms.
In this paper, we investigate the attractive properties of the proximal gradient algorithm with inertia. Notably, we show that using alternated inertia yields monotonically decreasing functional values, which contrasts with usual accelerated proximal gradient methods. We also provide convergence rates for the algorithm…
In this paper, we study volume growth, Liouville theorem and the local gradient estimate for -harmonic functions, and volume comparison property of unit balls in complete noncompact gradient Ricci shrinkers. We also study integral properties of f-harmonic functions and harmonic functions on such manifolds.
Paper studies asymmetric matrix sensing, proving gradient descent converges to low-rank solutions.
This paper studies gradient flows for sampling using various metrics and their affine invariance.
We analyze stochastic gradient descent for optimizing non-convex functions. In many cases for non-convex functions the goal is to find a reasonable local minimum, and the main concern is that gradient updates are trapped in saddle points. In this paper we identify strict saddle property for non-convex problem that allo…
Improved bounds for proximal gradient algorithms with computational errors.
Gradient descent variants improve phase retrieval accuracy.
Stochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated par…
SGD converges to an invariant distribution with sub-Gaussian or sub-exponential properties.
We study online convex optimization under stochastic sub-gradient observation faults, where we introduce adaptive algorithms with minimax optimal regret guarantees. We specifically study scenarios where our sub-gradient observations can be noisy or even completely missing in a stochastic manner. To this end, we propose…
The paper proves conditions for a manifold to have the Liouville property for the drifted Laplacian.
Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2…
This paper explores gradient flows for sampling distributions without normalization constants.
New optimization method combines gradient clipping and non-Euclidean smoothness.
Study on Einstein solitons with specific vector fields and their properties.
For a standard convolutional neural network, optimizing over the input pixels to maximize the score of some target class will generally produce a grainy-looking version of the original image. However, Santurkar et al. (2019) demonstrated that for adversarially-trained neural networks, this optimization produces images …
Researchers study heavy-tail properties of SGD using stochastic recurrence equations.
Bilevel optimization has been recently revisited for designing and analyzing algorithms in hyperparameter tuning and meta learning tasks. However, due to its nested structure, evaluating exact gradients for high-dimensional problems is computationally challenging. One heuristic to circumvent this difficulty is to use t…
Paper analyzes solutions to quasilinear elliptic equations on manifolds using Nash-Moser iteration.
Improves deep learning models by blending gradients from training loss and auxiliary objective.
In this paper, we propose a geometric framework to analyze the convergence properties of gradient descent trajectories in the context of linear neural networks. We translate a well-known empirical observation of linear neural nets into a conjecture that we call the \emph{overfitting conjecture} which states that, for a…
Gradient model for memristive systems in neurophysiology and neuromorphic circuits.
Most neural networks are trained using first-order optimization methods, which are sensitive to the parameterization of the model. Natural gradient descent is invariant to smooth reparameterizations because it is defined in a coordinate-free way, but tractable approximations are typically defined in terms of coordinate…
GeoAdaLer enhances geometric understanding of Adam for stochastic optimization.
Stochastic gradient descent procedures have gained popularity for parameter estimation from large data sets. However, their statistical properties are not well understood, in theory. And in practice, avoiding numerical instability requires careful tuning of key parameters. Here, we introduce implicit stochastic gradien…
Characterizes and examines gradient solitons on doubly warped product manifolds.
We analyze the learning properties of the stochastic gradient method when multiple passes over the data and mini-batches are allowed. We study how regularization properties are controlled by the step-size, the number of passes and the mini-batch size. In particular, we consider the square loss and show that for a unive…
Optimal transport distances are powerful tools to compare probability distributions and have found many applications in machine learning. Yet their algorithmic complexity prevents their direct use on large scale datasets. To overcome this challenge, practitioners compute these distances on minibatches {\em i.e.} they a…
In this paper, by slightly modifying Li-Yau's technique so that we can handle drifting Laplacians, we were able to find three different gradient estimates for the warping function, one for each sign of the Einstein constant of the fiber manifold. As an application, we exhibit a nonexistence theorem for gradient almost …
Investigates spectral properties of neural networks, showing invariance under certain conditions.
We investigate connections between information-theoretic and estimation-theoretic quantities in vector Poisson channel models. In particular, we generalize the gradient of mutual information with respect to key system parameters from the scalar to the vector Poisson channel model. We also propose, as another contributi…
Large deviations theory applied to policy gradient methods.
Study complete gradient Ricci solitons with zero radial Weyl curvature.
We derive lower bounds on the scalar curvature of complete non-compact gradient Yamabe solitons under some integral curvature conditions. Based on this, we prove that the corresponding potential functions have at most quadratic growth in distance. We also obtain a finite topological type property on complete shrinking …
Improved sampling method using regularized Stein Variational Gradient Flow.
Extends geometric structures to manifolds with new operators.
We study learning properties of accelerated gradient descent methods for linear least-squares in Hilbert spaces. We analyze the implicit regularization properties of Nesterov acceleration and a variant of heavy-ball in terms of corresponding learning error bounds. Our results show that acceleration can provides faster …
Proves properties of Morse vector fields on compact manifolds.
The study proves rotationally symmetric property of certain shrinking gradient Yamabe solitons.
Despite its empirical success and recent theoretical progress, there generally lacks a quantitative analysis of the effect of batch normalization (BN) on the convergence and stability of gradient descent. In this paper, we provide such an analysis on the simple problem of ordinary least squares (OLS). Since precise dyn…