Solves local minima problems on smooth manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
New metric shows how different regularization methods affect deep linear networks.
Nonlinear embedding manifold learning methods provide invaluable visual insights into the structure of high-dimensional data. However, due to a complicated nonconvex objective function, these methods can easily get stuck in local minima and their embedding quality can be poor. We propose a natural extension to several …
It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of sharpness/flatness are not invariant to rescaling of the network parameters, corresponding…
In this paper we shall illustrate that each polytopal moment-angle complex can be understood as the intersection of the minima of corresponding Siegel leaves and the unit sphere, with respect to the maximum norm. Consequently, an alternative proof of a rigidity theorem of Bosio and Meersseman is obtained; as piecewise …
Stochastic Gradient Descent (SGD) and its variants are mainstream methods for training deep networks in practice. SGD is known to find a flat minimum that often generalizes well. However, it is mathematically unclear how deep learning can select a flat minimum among so many minima. To answer the question quantitatively…
VAE global minima can learn correct manifold dimensions, even with conditioning variables.
Complex-valued neural networks avoid spurious local minima.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
On Riemannian manifolds of dimension 4, for prescribed scalar curvature equation, under lipschitzian condition on the prescribed curvature, we have an uniform estimate for the solutions of the equation if we control their minimas.
Truncated SGD with heavy-tailed noise eliminates sharp local minima.
Optimizers find approximate global minima in non-convex problems.
The study analyzes local minima in ReLU networks and finds low probability of bad local minima.
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…
Paper proposes faster method to find local minima in nonconvex optimization.
Global minima found for multidimensional scaling with penalties.
Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the core technique from these papers can be used to remove all bad local minima from any loss landscape, so long as the global minimum has a loss o…
Proposes NRS to find flat minima in deep neural networks.
Piecewise linear activations create many spurious local minima in neural networks.
SGD can jump from high rank minima to low rank minima in DLNs, but not back.
The notion of flat minima has played a key role in the generalization studies of deep learning models. However, existing definitions of the flatness are known to be sensitive to the rescaling of parameters. The issue suggests that the previous definitions of the flatness might not be a good measure of generalization, b…
Study reveals sharp characterisation of local minima in neural network loss landscapes.
It is well known that (stochastic) gradient descent has an implicit bias towards flat minima. In deep neural network training, this mechanism serves to screen out minima. However, the precise effect that this has on the trained network is not yet fully understood. In this paper, we characterize the flat minima in linea…
New insights into hidden minima in neural networks.
Study of SGD with state-dependent noise, improving escape from local minima.
We continue the comparison between lines of minima and Teichmueller geodesics begun in [CRS1]. We show that in the Teichmueller space of a surface S, lines of minima are quasi-geodesic with respect to the Teichmueller metric. The quasi-geodesic constants depend only on the topological type of S.
For one-hidden-layer ReLU networks, we prove that all differentiable local minima are global inside differentiable regions. We give the locations and losses of differentiable local minima, and show that these local minima can be isolated points or continuous hyperplanes, depending on an interplay between data, activati…
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer, or 2) at least as wide as the output layer. This result is the strongest possibl…
Recent advances in deep learning theory have evoked the study of generalizability across different local minima of deep neural networks (DNNs). While current work focused on either discovering properties of good local minima or developing regularization techniques to induce good local minima, no approach exists that ca…
Deep ReLU networks with extra parameters have mostly good loss landscapes.
In this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep n…
New findings suggest non-contrastive learning has many bad minima, not just collapsed ones.
Zeroth-order methods favor flat minima in machine learning.
We define lines of minima in the thick part of Outer space for the free group Fn with n>2 generators. We show that these lines of minima are contracting for the Lipschitz metric. Every fully irreducible outer automorphism of Fn defines such a line a minima. Now let G be a subgroup of the outer automorphism group of Fn …
New bounds link flat minima to good generalisation in overparameterized models.
Paper develops a new local convexity condition for non-isolated minima in non-convex optimization.
Black holes offer insights into machine learning's loss landscapes.
We investigate the dynamical and convergent properties of stochastic gradient descent (SGD) applied to Deep Neural Networks (DNNs). Characterizing the relation between learning rate, batch size and the properties of the final minima, such as width or generalization, remains an open question. In order to tackle this pro…
MCN improves deep neural networks by bettering local minima and generalizing well.
In this article we take up the calculation of the minimum number of colors needed to produce a non-trivial coloring of a knot. This is a knot invariant and we use the torus knots of type (2, n) as our case study. We calculate the minima in some cases. In other cases we estimate upper bounds for these minima leaning on …
DiMS sampler explores neural network loss minima via dissipative dynamics.
We present a novel optimization method, named the Combined Optimization Method (COM), for the joint optimization of two or more cost functions. Unlike the conventional joint optimization schemes, which try to find minima in a weighted sum of cost functions, the COM explores search space for common minima shared by all …
We consider the optimization problem associated with training simple ReLU neural networks of the form with respect to the squared loss. We provide a computer-assisted proof that even if the input distribution is standard Gaussian, even if the dime…
Let C be a real-analytic Jordan curve in . Then C cannot bound infinitely many disk-type minimal surfaces which provide relative minima of area.
New approach finds minima of geodesic lengths for non-uniform fillings.
Study minima of geodesic lengths for specific curves on surfaces.
Flat minima lead to better generalization in low-rank matrix recovery models.