Noise in RNNs promotes flatter minima and more stable dynamics.
problem Understanding and optimizing the training of RNNs with noise.
method Formalizing RNNs as stochastic differential equations and analyzing the effect of noise in the hidden states.
result Noise injection in RNNs leads to flatter minima, more stable dynamics, and improved robustness.
Paper studies spaces of flattenings of simplicial spheres and their homotopy type.
problem Existence and uniqueness of differentiable structures on simplicial spheres.
method Analyzes spaces of flattenings of simplicial spheres and their homotopy type.
result Spaces of flattenings have the homotopy type of the orthogonal group.
Flattenings of knotted surfaces help define new invariants.
problem Understanding and quantifying knotted surfaces in 4-sphere.
method Using hyperbolic decompositions and projections onto 2-sphere.
result Introduced layering, trunk, and partition number invariants.
This study develops a NURBS-based method for conformal surface flattening without singularities.
problem Flatten surfaces conformally without singularities.
method NURBS-based approach with iterative refinement of input and flattening surfaces, leveraging nonlinear extension of VarPro.
result Developed a singularity-free NURBS-based method for conformal surface flattening.
New neural networks flatten and reconstruct manifolds from samples.
problem Learning from high-dimensional data embedded in submanifolds.
method Flattening Networks (FlatNet) that linearize and reconstruct embedded submanifolds.
result FlatNet achieves balance of interpretability, feasibility, and generalization.
We discuss Ghys' theorem on 4 zeroes of the Schwarzian derivative and its relation with flattening points of Legendrian curves and Sturm theory.
This paper continues the previous studies in two papers of Huang-Yin [HY3-4] on the flattening problem of a CR singular point of real codimension two sitting in a submanifold in Cn+1 with n+1≥3, whose CR points are non-minimal. Partially based on the geometric approach initiated in [HY3] and a forma…
A primary goal in this paper is to study the question that asks when a real analytic submanifold M in Cn+1 bounds a real analytic (up to M) Levi-flat hypersurface M^ near p∈M such that M^ is foliated by a family of complex hypersurfaces moving along the normal direction of M at …
In this paper, we discuss centroaffine geometry of polygons in 3-space. For a polygon X that is locally convex with respect to an origin together with a transversal vector field U, we define the centroaffine dual pair (Y,V) similarly to [6]. We prove that vertices of (X,U) correspond to flattening points for …
In this paper, we are concerned with the problem of creating flattening maps of simply-connected open surfaces in R3. Using a natural principle of density diffusion in physics, we propose an effective algorithm for computing density-equalizing flattening maps with any prescribed density distribution. By var…
Any smooth surface in R^3 may be flattened along the z-axis, and the flattened surface becomes close to a billiard table in R^2 . We show that, under some hypotheses, the geodesic flow of this surface converges locally uniformly to the billiard flow. Moreover, if the billiard is dispersive and has finite horizon, then …
SGD favors flat minima exponentially more than sharp minima in deep learning.
problem Understanding how SGD selects flat minima in deep learning.
method Developed a density diffusion theory (DDT) to analyze minima selection.
result SGD exponentially favors flat minima over sharp minima due to Hessian-dependent noise.
AWP improves robustness by flattening weight loss landscape.
problem Improving robustness of deep neural networks against adversarial examples.
method Explicitly regularizes the flatness of weight loss landscape through adversarial weight perturbation.
result AWP forms a double-perturbation mechanism in adversarial training, leading to flatter weight loss landscape.
Complex-valued neural networks avoid spurious local minima.
problem Finding spurious local minima in neural networks.
method Proved no spurious local minima for shallow complex neural networks with quadratic activations.
result Complex-valued weights eliminate spurious local minima in neural networks.
We consider in this paper the FRS-deformations of a family of space curves with codimension ≤3. Some geometric aspects of a space curve such as flattenings, vertices and twistings points has been studied.
Study reveals properties of local minima in ReLU networks.
problem Understanding the loss landscape of neural networks.
method Theoretical analysis of one-hidden-layer ReLU networks.
result All differentiable local minima are global within certain regions.
The study proves a discrete version of Segre's theorem for polygonal curves.
problem Proving a discrete analog of a four-vertex theorem for spherical curves.
method Using the concept of discrete tangent indicatrix of a polygon.
result A polygon with at least four vertices and a non-self-intersecting discrete tangent indicatrix has at least four flattenings.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.
Theory explains power-law distributions without complex models.
problem Understanding power-law distributions in geometrically growing systems.
method Developed a theory of geometrically growing systems and applied it to explain various distributions.
result The geometrically growing system's distribution flattens over time, increasing relative size ratios.
Method flattens complex surfaces with consistent density and shape.
problem Shape deformations and local geometric distortions in density-equalizing maps for multiply-connected surfaces.
method Formulates density diffusion as a quasiconformal flow, solving an energy minimization problem involving the Beltrami coefficient to ensure bijectivity and control distortion.
result Achieves optimal parameterization of multiply-connected surfaces with bijective and controlled geometric distortions.
Gradient descent in deep networks tends to find flat minima, which are nearly balanced.
problem Understanding the effect of gradient descent on the structure of minima in deep neural networks.
method Characterized flat minima in linear neural networks trained with a quadratic loss.
result Flat minima correspond to nearly balanced networks where the gain from input to intermediate representations is nearly constant.
Truncated SGD with heavy-tailed noise eliminates sharp local minima.
problem Avoiding sharp local minima in deep learning models.
method Truncated SGD with heavy-tailed gradient noise.
result Truncated SGD can eliminate sharp local minima entirely from its training trajectory.
We experimentally achieve a 19% capacity gain per Watt of electrical supply power in a 12-span link by eliminating gain flattening filters and optimizing launch powers using machine learning by deep neural networks in a massively parallel fiber context.
Optimizers find approximate global minima in non-convex problems.
problem Understanding why local methods solve non-convex optimization problems.
method Formalizing the hypothesis that many local minima are approximately global minima.
result Most local minima of practical non-convex objectives are approximately global minima.
A mathematical model describes deforming manifolds with precise vectors and fields.
problem Modeling and describing the deformation of complex manifolds in practical applications.
method Proposes a modified differential dynamic model with constraints on spatial and temporal continuity, presenting deforming vector and field.
result Demonstrates the effectiveness of an autonomous deforming field in data dimension reduction tasks.
The study analyzes local minima in ReLU networks and finds low probability of bad local minima.
problem Understanding the existence and probability of local minima in ReLU networks.
method Theoretical analysis combined with linear programming and experiments on MNIST and CIFAR-10 datasets.
result No bad differentiable local minima found almost everywhere in weight space.
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…
A method to automatically and symbolically detect and resolve degenerate parameter combinations from parameter-data pairs.
problem Identifying degenerate parameter combinations in physical models or real-world datasets.
method The degeneracy distillery method detects and resolves degenerate parameter combinations from parameter-data pairs.
result The method reduces the simulation budget required for downstream neural posterior estimation.
Paper proposes faster method to find local minima in nonconvex optimization.
problem Escaping saddle points and finding local minima in nonconvex optimization.
method LENA (Last stEp shriNkAge) framework for faster perturbed stochastic gradient methods.
result LENA finds (ε,εH)-approximate local minima within ildeO(ε−3+εH−6) evaluations. Global minima found for multidimensional scaling with penalties.
problem Finding global minima in multidimensional scaling.
method Combining stress loss function with a quadratic penalty term to find minimizers.
result Trajectory of minimizers leads to global minima.
Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the core technique from these papers can be used to remove all bad local minima from any loss landscape, so long as the global minimum has a loss o…
Proposes NRS to find flat minima in deep neural networks.
problem Finding optimal solutions in deep neural networks with overparameterization.
method NRS leverages the concept of flat minima and uses Kullback-Leibler divergence to regularize the neighborhood region in weight space.
result NRS drives optimizers towards flat minima, improving generalization ability across various model architectures.
Piecewise linear activations create many spurious local minima in neural networks.
problem Understanding the loss surface of neural networks with piecewise linear activations.
method Proved the existence of infinite spurious local minima and partitioned the loss surface into smooth cells.
result Piecewise linear activations create many spurious local minima that are invariant under a continuous path.
SGD can jump from high rank minima to low rank minima in DLNs, but not back.
problem SGD's tendency to get stuck in high rank minima in DLNs.
method Analysis of the L2-regularized loss function of DLNs and the definition of absorbing sets. result SGD has a non-zero probability to jump from high rank minima to low rank minima but zero probability to jump back.
The notion of flat minima has played a key role in the generalization studies of deep learning models. However, existing definitions of the flatness are known to be sensitive to the rescaling of parameters. The issue suggests that the previous definitions of the flatness might not be a good measure of generalization, b…
Study reveals sharp characterisation of local minima in neural network loss landscapes.
problem Characterizing local minima in high-dimensional two-layer ReLU neural networks.
method Exact low-dimensional representation of local minima using summary statistics and link with one-pass SGD dynamics.
result Local minima in overparameterized neural networks form discrete families with varying stability and reachability.
New insights into hidden minima in neural networks.
problem Identifying hidden minima in two-layer ReLU networks.
method Analyzing curves along which loss is minimized, focusing on eigenvalue contributions.
result Distinctive structural and symmetry properties of arcs emanating from hidden minima.
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
Study of SGD with state-dependent noise, improving escape from local minima.
problem Understanding and improving the dynamics of SGD in non-convex optimization.
method Formal study on SGD with state-dependent noise, proposing power-law dynamic with state-dependent diffusion.
result Power-law dynamic can escape from sharp minima exponentially faster than flat minima.
We continue the comparison between lines of minima and Teichmueller geodesics begun in [CRS1]. We show that in the Teichmueller space of a surface S, lines of minima are quasi-geodesic with respect to the Teichmueller metric. The quasi-geodesic constants depend only on the topological type of S.
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer, or 2) at least as wide as the output layer. This result is the strongest possibl…
A single-vertex origami is a piece of paper with straight-line rays called creases emanating from a fold vertex placed in its interior or on its boundary. The Single-Vertex Origami Flattening problem asks whether it is always possible to reconfigure the creased paper from any configuration compatible with the metric, t…
Recent advances in deep learning theory have evoked the study of generalizability across different local minima of deep neural networks (DNNs). While current work focused on either discovering properties of good local minima or developing regularization techniques to induce good local minima, no approach exists that ca…
A new algorithm flattens multi-modal distributions for better deep learning.
problem Bayesian learning in big data with multi-modal distributions.
method Contour Stochastic Gradient Langevin Dynamics (CSGLD) algorithm.
result The CSGLD algorithm avoids local traps in deep neural networks.
Deep ReLU networks with extra parameters have mostly good loss landscapes.
problem Finding good local minima in the loss landscape of deep neural networks.
method Analyzing shallow and deep ReLU networks with extra parameters on a generic dataset.
result Most activation patterns correspond to regions with no bad local minima.
In this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep n…
New findings suggest non-contrastive learning has many bad minima, not just collapsed ones.
problem The effectiveness of non-contrastive learning in unsupervised feature learning.
method Theoretical analysis and controlled experiments on simple data models.
result Non-contrastive losses have a preponderance of non-collapsed bad minima, and these minima are not avoided during training.
Zeroth-order methods favor flat minima in machine learning.
problem Finding solutions with small Hessian trace in optimization.
method Zeroth-order optimization with two-point estimator.
result Zeroth-order optimization converges to flat minima.