A new pruning method reduces neural network computation without retraining.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD). Recent work conjectures that SGD with proper configuration is able to find wide and flat local minima, which have been proposed to be associated with good gener…
A new pruning method finds sparse minimizers in flat regions of deep neural networks.
Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenval…
Quantization-aware training can recover accuracy lost by post-training quantization.
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large…
A D-Wave quantum annealer (QA) having a 2048 qubit lattice, with no missing qubits and couplings, allowed embedding of a complete graph of a Restricted Boltzmann Machine (RBM). A handwritten digit OptDigits data set having 8x7 pixels of visible units was used to train the RBM using a classical Contrastive Divergence. E…
A new Kolmogorov-Arnold network improves function approximation and optimization.
A common difficulty in applications of machine learning is the lack of any general principle for guiding the choices of key parameters of the underlying neural network. Focusing on a class of recurrent neural networks - reservoir computing systems that have recently been exploited for model-free prediction of nonlinear…
It was empirically confirmed by Keskar et al.\cite{SharpMinima} that flatter minima generalize better. However, for the popular ReLU network, sharp minimum can also generalize well \cite{SharpMinimacan}. The conclusion demonstrates that the existing definitions of flatness fail to account for the complex geometry of Re…
We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that from any point in parameter space there exists a continuous path on which the cross-entropy loss is non-increasing and gets arbitrarily clos…
In this article we apply a Bochner type formula to show that on a compact conformally flat riemannian manifold (or half-conformally flat in dimension 4) certain types of orthogonal almost-complex structures, if they exist, give the absolute minimum for the energy functional. We give a few examples when such minimizers …
This paper examines SVB's failure and its impact on bank stocks.
SALR improves deep learning generalization by dynamically adjusting learning rates.
Study minimum ribbonlength of immersed flat knots and links.
A new AI optimization method uses energy-conserving dynamics inspired by Born-Infeld theory.
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima. In a network of hidden layers with neurons in layers $k = 1, \ldots, …
SAM selects flatter minima late in training, improving generalization.
Proposes a method to partition univariate data into unimodal subsets.
Let be a compact Ricci-flat 4-manifold. For let (respectively ) denote the maximum (respectively the minimum) of sectional curvatures at . We prove that if for all , for some constant with , th…
If X is a full, finitely generated, projective module over a non-commutative torus, the Yang-Mills functional attains its minimum exactly on the flat connections on X. We classify the flat connections on modules admitting integrable connections.
New framework reveals thermodynamic principles for LLM training.
We theoretically study the landscape of the training error for neural networks in overparameterized cases. We consider three basic methods for embedding a network into a wider one with more hidden units, and discuss whether a minimum point of the narrower network gives a minimum or saddle point of the wider one. Our re…
SGD favors flat minima exponentially more than sharp minima in deep learning.
We solve the optimization of two-layer ReLU networks using convex math.
DiMS sampler explores neural network loss minima via dissipative dynamics.
WSD schedule improves model training efficiency by adapting learning rates dynamically.
Flat plumbing basket surfaces of links were introduced to study the geometry of the complement of the links. These flat plumbing basket surface can be presented by a sequential presentation known as flat plumbing basket code first found by Furihata, Hirasawa and Kobayashi. The minimum number of flat plumbings to obtain…
Every link is shown to be presentable as a boundary of an unknotted flat banded surface. A (flat) banded link is defined as a boundary of an unknotted (flat) banded surface. A link's (flat) band index is defined as the minimum number of bands required to present the link as boundaries of an unknotted (flat) banded surf…
HyPV-LEAD detects cryptocurrency anomalies proactively, improving financial security.
This paper presents an algorithm for a complete and efficient calibration of the Heston stochastic volatility model. We express the calibration as a nonlinear least squares problem. We exploit a suitable representation of the Heston characteristic function and modify it to avoid discontinuities caused by branch switchi…
The performance of deep neural networks is often attributed to their automated, task-related feature construction. It remains an open question, though, why this leads to solutions with good generalization, even in cases where the number of parameters is larger than the number of samples. Back in the 90s, Hochreiter and…
Study on folded ribbon knots and their minimum length.
Combining insights from machine learning and quantum Monte Carlo, the stochastic reconfiguration method with neural network Ansatz states is a promising new direction for high-precision ground state estimation of quantum many-body problems. Even though this method works well in practice, little is known about the learn…
Continuous-time analysis shows SGD with noise prefers flat minima.
Solves local minima problems on smooth manifolds.
We construct a Fourier--Mukai transform for smooth complex vector bundles over a torus bundle the vector bundles being endowed with various structures of increasing complexity. At a minimum, we consider vector bundles with a flat partial unitary connection, that is families or deformations of flat …
Unbalanced data arises in many learning tasks such as clustering of multi-class data, hierarchical divisive clustering and semisupervised learning. Graph-based approaches are popular tools for these problems. Graph construction is an important aspect of graph-based learning. We show that graph-based algorithms can fail…
Local gap theorem for Ricci shrinkers ensures flatness if certain functionals are close to zero.
Warm starts improve variational quantum algorithms by avoiding barren plateaus.
SGD noise helps select flat minima by concentrating in sharp directions and being proportional to loss value.
Designs chiral photonic structures using machine learning for efficient optical properties.