Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Nov 199319922001200920172026
48 results for Minima Flattening

Noise in RNNs promotes flatter minima and more stable dynamics.

problem Understanding and optimizing the training of RNNs with noise.
method Formalizing RNNs as stochastic differential equations and analyzing the effect of noise in the hidden states.
result Noise injection in RNNs leads to flatter minima, more stable dynamics, and improved robustness.

This study develops a NURBS-based method for conformal surface flattening without singularities.

problem Flatten surfaces conformally without singularities.
method NURBS-based approach with iterative refinement of input and flattening surfaces, leveraging nonlinear extension of VarPro.
result Developed a singularity-free NURBS-based method for conformal surface flattening.

This paper continues the previous studies in two papers of Huang-Yin [HY3-4] on the flattening problem of a CR singular point of real codimension two sitting in a submanifold in Cn+1{\mathbb C}^{n+1} with n+13n+1\ge 3, whose CR points are non-minimal. Partially based on the geometric approach initiated in [HY3] and a forma…

2017-03-27abs ↗pdf ↗

In this paper, we discuss centroaffine geometry of polygons in 33-space. For a polygon XX that is locally convex with respect to an origin together with a transversal vector field UU, we define the centroaffine dual pair (Y,V)(Y,V) similarly to [6]. We prove that vertices of (X,U)(X,U) correspond to flattening points for …

2018-12-03abs ↗pdf ↗

In this paper, we are concerned with the problem of creating flattening maps of simply-connected open surfaces in R3\mathbb{R}^3. Using a natural principle of density diffusion in physics, we propose an effective algorithm for computing density-equalizing flattening maps with any prescribed density distribution. By var…

2017-04-08abs ↗pdf ↗

Any smooth surface in R^3 may be flattened along the z-axis, and the flattened surface becomes close to a billiard table in R^2 . We show that, under some hypotheses, the geodesic flow of this surface converges locally uniformly to the billiard flow. Moreover, if the billiard is dispersive and has finite horizon, then …

2015-03-14abs ↗pdf ↗

SGD favors flat minima exponentially more than sharp minima in deep learning.

problem Understanding how SGD selects flat minima in deep learning.
method Developed a density diffusion theory (DDT) to analyze minima selection.
result SGD exponentially favors flat minima over sharp minima due to Hessian-dependent noise.

AWP improves robustness by flattening weight loss landscape.

problem Improving robustness of deep neural networks against adversarial examples.
method Explicitly regularizes the flatness of weight loss landscape through adversarial weight perturbation.
result AWP forms a double-perturbation mechanism in adversarial training, leading to flatter weight loss landscape.

The study proves a discrete version of Segre's theorem for polygonal curves.

problem Proving a discrete analog of a four-vertex theorem for spherical curves.
method Using the concept of discrete tangent indicatrix of a polygon.
result A polygon with at least four vertices and a non-self-intersecting discrete tangent indicatrix has at least four flattenings.

Unified approach to characterize and regularize deep neural network local minima.

problem Characterize and improve generalizability of deep neural network local minima.
method Information-theoretic Fisher information metric for local minima characterization and regularization.
result Unified approach successfully characterizes and improves generalizability of DNNs.

Paper finds wide minima are better for generalization and proposes a new learning rate schedule.

problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.

Method flattens complex surfaces with consistent density and shape.

problem Shape deformations and local geometric distortions in density-equalizing maps for multiply-connected surfaces.
method Formulates density diffusion as a quasiconformal flow, solving an energy minimization problem involving the Beltrami coefficient to ensure bijectivity and control distortion.
result Achieves optimal parameterization of multiply-connected surfaces with bijective and controlled geometric distortions.

Theory explains power-law distributions without complex models.

problem Understanding power-law distributions in geometrically growing systems.
method Developed a theory of geometrically growing systems and applied it to explain various distributions.
result The geometrically growing system's distribution flattens over time, increasing relative size ratios.

Gradient descent in deep networks tends to find flat minima, which are nearly balanced.

problem Understanding the effect of gradient descent on the structure of minima in deep neural networks.
method Characterized flat minima in linear neural networks trained with a quadratic loss.
result Flat minima correspond to nearly balanced networks where the gain from input to intermediate representations is nearly constant.

A mathematical model describes deforming manifolds with precise vectors and fields.

problem Modeling and describing the deformation of complex manifolds in practical applications.
method Proposes a modified differential dynamic model with constraints on spatial and temporal continuity, presenting deforming vector and field.
result Demonstrates the effectiveness of an autonomous deforming field in data dimension reduction tasks.

The study analyzes local minima in ReLU networks and finds low probability of bad local minima.

problem Understanding the existence and probability of local minima in ReLU networks.
method Theoretical analysis combined with linear programming and experiments on MNIST and CIFAR-10 datasets.
result No bad differentiable local minima found almost everywhere in weight space.

In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…

2017-02-27abs ↗pdf ↗

A method to automatically and symbolically detect and resolve degenerate parameter combinations from parameter-data pairs.

problem Identifying degenerate parameter combinations in physical models or real-world datasets.
method The degeneracy distillery method detects and resolves degenerate parameter combinations from parameter-data pairs.
result The method reduces the simulation budget required for downstream neural posterior estimation.

Paper proposes faster method to find local minima in nonconvex optimization.

problem Escaping saddle points and finding local minima in nonconvex optimization.
method LENA (Last stEp shriNkAge) framework for faster perturbed stochastic gradient methods.
result LENA finds (ε,εH)(ε, ε_{H})-approximate local minima within ildeO(ε3+εH6) ilde O(ε^{-3} + ε_{H}^{-6}) evaluations.

Piecewise linear activations create many spurious local minima in neural networks.

problem Understanding the loss surface of neural networks with piecewise linear activations.
method Proved the existence of infinite spurious local minima and partitioned the loss surface into smooth cells.
result Piecewise linear activations create many spurious local minima that are invariant under a continuous path.

Proposes NRS to find flat minima in deep neural networks.

problem Finding optimal solutions in deep neural networks with overparameterization.
method NRS leverages the concept of flat minima and uses Kullback-Leibler divergence to regularize the neighborhood region in weight space.
result NRS drives optimizers towards flat minima, improving generalization ability across various model architectures.

SGD can jump from high rank minima to low rank minima in DLNs, but not back.

problem SGD's tendency to get stuck in high rank minima in DLNs.
method Analysis of the L2L_{2}-regularized loss function of DLNs and the definition of absorbing sets.
result SGD has a non-zero probability to jump from high rank minima to low rank minima but zero probability to jump back.

Study reveals sharp characterisation of local minima in neural network loss landscapes.

problem Characterizing local minima in high-dimensional two-layer ReLU neural networks.
method Exact low-dimensional representation of local minima using summary statistics and link with one-pass SGD dynamics.
result Local minima in overparameterized neural networks form discrete families with varying stability and reachability.

In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…

2010-08-30abs ↗pdf ↗

Study of SGD with state-dependent noise, improving escape from local minima.

problem Understanding and improving the dynamics of SGD in non-convex optimization.
method Formal study on SGD with state-dependent noise, proposing power-law dynamic with state-dependent diffusion.
result Power-law dynamic can escape from sharp minima exponentially faster than flat minima.

We continue the comparison between lines of minima and Teichmueller geodesics begun in [CRS1]. We show that in the Teichmueller space of a surface S, lines of minima are quasi-geodesic with respect to the Teichmueller metric. The quasi-geodesic constants depend only on the topological type of S.

2007-06-14abs ↗pdf ↗

A single-vertex origami is a piece of paper with straight-line rays called creases emanating from a fold vertex placed in its interior or on its boundary. The Single-Vertex Origami Flattening problem asks whether it is always possible to reconfigure the creased paper from any configuration compatible with the metric, t…

2010-03-17abs ↗pdf ↗

A new algorithm flattens multi-modal distributions for better deep learning.

problem Bayesian learning in big data with multi-modal distributions.
method Contour Stochastic Gradient Langevin Dynamics (CSGLD) algorithm.
result The CSGLD algorithm avoids local traps in deep neural networks.

In this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep n…

2019-01-02abs ↗pdf ↗

New findings suggest non-contrastive learning has many bad minima, not just collapsed ones.

problem The effectiveness of non-contrastive learning in unsupervised feature learning.
method Theoretical analysis and controlled experiments on simple data models.
result Non-contrastive losses have a preponderance of non-collapsed bad minima, and these minima are not avoided during training.