Paper develops a new local convexity condition for non-isolated minima in non-convex optimization.
problem Lack of theory for non-isolated minima in non-convex optimization.
method Formulates a new local convexity condition and studies SGD convergence under this condition.
result Shows SGD can converge locally under the new condition.
Formula for sections on complex manifolds with non-isolated components.
problem Localization of sections on complex manifolds with non-isolated zero varieties.
method Logarithmic Bott localization formula, current-theoretic formulation.
result Established a formula for sections on compact complex manifolds with non-isolated components.
Defines Perelman's functionals on manifolds with non-isolated conical singularities.
problem Defining functionals on manifolds with non-isolated conical singularities.
method Starting from a spectral point of view for the Perelman's λ-functional, defining the spectrum of Schrödinger operator and proving the existence of discrete eigenvalues.
result Proves the existence of the infimum of W-functional and obtains asymptotic behavior of eigenfunctions.
For two complex vector bundles admitting a homomorphism between them, a Poincaré-Hopf formula for the difference of the Chern character numbers of these two vector bundles with isolated singularities is established by Huitao Feng, Weiping Li and Weiping Zhang. This article extend their reslut about Poincaré-Hopf type f…
Study of non-convex potential functions in deep learning with Poincaré inequality.
problem Understanding convergence of stochastic dynamics in non-convex potential landscapes.
method Introduced log-Polyak-Lojasiewicz (log-PL) measures and analyzed their convergence properties.
result Langevin dynamics converges at a rate of O~(1/ε) for sufficiently small ε. Khimshiashvili proved a topological degree formula for the Eu-ler characteristic of the Milnor fibres of a real function-germ with an isolated singularity. We give two generalizations of this result for non-isolated singularities. As corollaries we obtain an algebraic formula for the Euler characteristic of the fibres …
Defines Milnor number for foliations and shows its topological invariance.
problem Defining and proving invariance of Milnor number for non-isolated singularities of holomorphic foliations.
method Defining Milnor number as intersection number of sections; proving invariance via C1 topological equivalences. result Milnor number is invariant under C1 topological equivalences. We generalize the Novikov inequalities for 1-forms in two different directions: first, we allow non-isolated critical points (assuming that they are non-degenerate in the sense of R.Bott), and, secondly, we strengthen the inequalities by means of twisting by an arbitrary flat bundle. The proof uses Bismut's modificatio…
Develops a formalism for studying general horizons and derives a near-horizon equation.
problem Analyzes the geometry of general horizons in spacetime.
method Introduces a formalism based on encoding the zeroth and first transverse derivatives of the deformation tensor on null hypersurfaces.
result Derives a generalized near-horizon equation that holds on any horizon.
We study the boundary L_t of the Milnor fiber for the non-isolated singularities in C^3 with equation z^m - g(x,y) = 0 where g(x,y) is a non-reduced plane curve germ. We give a complete proof that L_t is a Waldhausen graph manifold and we provide the tools to construct its plumbing graph. As an example, we give the plu…
We define and study the secondary Chern-Euler class for a general submanifold of a Riemannian manifold. Using this class, we define and study index for a vector field with non-isolated singularities on a submanifold. As an application, our studies give conceptual proofs of a classical result of Chern.
Extract anomalies from 5D SCFTs using extra-dimensional η-invariants.
problem Anomalies in quantum field theories.
method Use extra-dimensional η-invariants to bypass traditional blowup techniques.
result Anomalies can be determined directly from η-invariants of asymptotic boundaries.
Solutions to scalar curvature equations have the property that all possible blow-up points are isolated, at least in low dimensions. This property is commonly used as the first step in the proofs of compactness. We show that this result becomes false for some arbitrarily small, smooth perturbations of the potential.
SGD favors flat minima exponentially more than sharp minima in deep learning.
problem Understanding how SGD selects flat minima in deep learning.
method Developed a density diffusion theory (DDT) to analyze minima selection.
result SGD exponentially favors flat minima over sharp minima due to Hessian-dependent noise.
The paper constructs instanton complexes on stratified pseudomanifolds.
problem Analyzing functions with non-isolated critical points on singular spaces.
method Constructing Witten instanton complexes and Hilbert complexes.
result Proves Morse inequalities for stratified pseudomanifolds.
Complex-valued neural networks avoid spurious local minima.
problem Finding spurious local minima in neural networks.
method Proved no spurious local minima for shallow complex neural networks with quadratic activations.
result Complex-valued weights eliminate spurious local minima in neural networks.
Study characteristic classes of a specific type of determinantal varieties.
problem Understanding the geometric properties of a special class of determinantal varieties.
method Used Schubert calculus to derive explicit formulas for Chern-Schwartz-MacPherson and Chern-Mather classes.
result Explicit formulas for sectional Euler characteristics, characteristic cycles, and polar classes were obtained.
Study reveals properties of local minima in ReLU networks.
problem Understanding the loss landscape of neural networks.
method Theoretical analysis of one-hidden-layer ReLU networks.
result All differentiable local minima are global within certain regions.
The paper calculates critical configurations and Morse indices for polygons on circles or ellipses.
problem Finding critical configurations and their properties for polygons on circles or ellipses.
method Computing Morse indices and gradient vector fields for isolated critical points, relating to eigenvalue questions.
result Computed Morse indices and relationships to eigenvalue questions for polygons on circles or ellipses.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.
We prove the equivariant holomorphic Morse inequalities for a holomorphic torus action on a holomorphic vector bundle over a compact Kahler manifold when the fixed-point set is not necessarily discrete. Such inequalities bound the twisted Dolbeault cohomologies of the Kahler manifold in terms of those of the fixed-poin…
Gradient descent in deep networks tends to find flat minima, which are nearly balanced.
problem Understanding the effect of gradient descent on the structure of minima in deep neural networks.
method Characterized flat minima in linear neural networks trained with a quadratic loss.
result Flat minima correspond to nearly balanced networks where the gain from input to intermediate representations is nearly constant.
Truncated SGD with heavy-tailed noise eliminates sharp local minima.
problem Avoiding sharp local minima in deep learning models.
method Truncated SGD with heavy-tailed gradient noise.
result Truncated SGD can eliminate sharp local minima entirely from its training trajectory.
Study stabilizers of smooth functions on surfaces, focusing on Morse-Bott functions.
problem Understanding the homotopy type of stabilizers of smooth functions on surfaces.
method Analyzing the homotopy properties of stabilizers for a specific class of smooth functions.
result The homotopy type of the connected component of the identity map of the stabilizer is completely described for Morse-Bott functions.
Optimizers find approximate global minima in non-convex problems.
problem Understanding why local methods solve non-convex optimization problems.
method Formalizing the hypothesis that many local minima are approximately global minima.
result Most local minima of practical non-convex objectives are approximately global minima.
The study analyzes local minima in ReLU networks and finds low probability of bad local minima.
problem Understanding the existence and probability of local minima in ReLU networks.
method Theoretical analysis combined with linear programming and experiments on MNIST and CIFAR-10 datasets.
result No bad differentiable local minima found almost everywhere in weight space.
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…
We generalize the Novikov inequalities for 1-forms in two different directions: first, we allow non-isolated critical points (assuming that they are non-degenerate in the sense of R.Bott), and, secondly, we strengthen the inequalities by means of twisting by an arbitrary flat bundle. We also obtain an L2 version of …
Paper proposes faster method to find local minima in nonconvex optimization.
problem Escaping saddle points and finding local minima in nonconvex optimization.
method LENA (Last stEp shriNkAge) framework for faster perturbed stochastic gradient methods.
result LENA finds (ε,εH)-approximate local minima within ildeO(ε−3+εH−6) evaluations. Global minima found for multidimensional scaling with penalties.
problem Finding global minima in multidimensional scaling.
method Combining stress loss function with a quadratic penalty term to find minimizers.
result Trajectory of minimizers leads to global minima.
Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the core technique from these papers can be used to remove all bad local minima from any loss landscape, so long as the global minimum has a loss o…
Proposes NRS to find flat minima in deep neural networks.
problem Finding optimal solutions in deep neural networks with overparameterization.
method NRS leverages the concept of flat minima and uses Kullback-Leibler divergence to regularize the neighborhood region in weight space.
result NRS drives optimizers towards flat minima, improving generalization ability across various model architectures.
Piecewise linear activations create many spurious local minima in neural networks.
problem Understanding the loss surface of neural networks with piecewise linear activations.
method Proved the existence of infinite spurious local minima and partitioned the loss surface into smooth cells.
result Piecewise linear activations create many spurious local minima that are invariant under a continuous path.
SGD can jump from high rank minima to low rank minima in DLNs, but not back.
problem SGD's tendency to get stuck in high rank minima in DLNs.
method Analysis of the L2-regularized loss function of DLNs and the definition of absorbing sets. result SGD has a non-zero probability to jump from high rank minima to low rank minima but zero probability to jump back.
The notion of flat minima has played a key role in the generalization studies of deep learning models. However, existing definitions of the flatness are known to be sensitive to the rescaling of parameters. The issue suggests that the previous definitions of the flatness might not be a good measure of generalization, b…
Study reveals sharp characterisation of local minima in neural network loss landscapes.
problem Characterizing local minima in high-dimensional two-layer ReLU neural networks.
method Exact low-dimensional representation of local minima using summary statistics and link with one-pass SGD dynamics.
result Local minima in overparameterized neural networks form discrete families with varying stability and reachability.
New insights into hidden minima in neural networks.
problem Identifying hidden minima in two-layer ReLU networks.
method Analyzing curves along which loss is minimized, focusing on eigenvalue contributions.
result Distinctive structural and symmetry properties of arcs emanating from hidden minima.
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
Study of SGD with state-dependent noise, improving escape from local minima.
problem Understanding and improving the dynamics of SGD in non-convex optimization.
method Formal study on SGD with state-dependent noise, proposing power-law dynamic with state-dependent diffusion.
result Power-law dynamic can escape from sharp minima exponentially faster than flat minima.
We continue the comparison between lines of minima and Teichmueller geodesics begun in [CRS1]. We show that in the Teichmueller space of a surface S, lines of minima are quasi-geodesic with respect to the Teichmueller metric. The quasi-geodesic constants depend only on the topological type of S.
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer, or 2) at least as wide as the output layer. This result is the strongest possibl…
Recent advances in deep learning theory have evoked the study of generalizability across different local minima of deep neural networks (DNNs). While current work focused on either discovering properties of good local minima or developing regularization techniques to induce good local minima, no approach exists that ca…
Deep ReLU networks with extra parameters have mostly good loss landscapes.
problem Finding good local minima in the loss landscape of deep neural networks.
method Analyzing shallow and deep ReLU networks with extra parameters on a generic dataset.
result Most activation patterns correspond to regions with no bad local minima.
In this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep n…
New findings suggest non-contrastive learning has many bad minima, not just collapsed ones.
problem The effectiveness of non-contrastive learning in unsupervised feature learning.
method Theoretical analysis and controlled experiments on simple data models.
result Non-contrastive losses have a preponderance of non-collapsed bad minima, and these minima are not avoided during training.
Zeroth-order methods favor flat minima in machine learning.
problem Finding solutions with small Hessian trace in optimization.
method Zeroth-order optimization with two-point estimator.
result Zeroth-order optimization converges to flat minima.
Solves local minima problems on smooth manifolds.
problem Local minima issues on smooth manifolds.
method Introducing valley functions and applying Morse's lemma.
result Eliminates critical points and reduces to 1D.
We define lines of minima in the thick part of Outer space for the free group Fn with n>2 generators. We show that these lines of minima are contracting for the Lipschitz metric. Every fully irreducible outer automorphism of Fn defines such a line a minima. Now let G be a subgroup of the outer automorphism group of Fn …