We propose to study neural networks' loss surfaces by methods of topological data analysis. We suggest to apply barcodes of Morse complexes to explore topology of loss surfaces. An algorithm for calculations of the loss function's barcodes of local minima is described. We have conducted experiments for calculating barc…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Random Matrix Theory explains loss surface Hessians in neural networks.
Quantification of the stationary points and the associated basins of attraction of neural network loss surfaces is an important step towards a better understanding of neural network loss surfaces at large. This work proposes a novel method to visualise basins of attraction together with the associated stationary points…
New methods connect low-loss points on neural network surfaces.
Paper discovers simplicial complexes connecting trained models for improved ensembling.
New tools reveal simple structure in complex hyperparameter loss surfaces near optima.
We describe loss surfaces using topological Betti numbers.
Piecewise linear activations create many spurious local minima in neural networks.
Study compact Willmore surfaces without complex structure convergence, computing energy loss and geodesic lengths.
It has been argued in the past that high-dimensional neural networks do not exhibit local minima capable of trapping an optimisation algorithm. However, the relationship between loss surface modality and the neural architecture parameters, such as the number of hidden neurons per layer and the number of hidden layers, …
Explains deep learning models and their geometric properties.
The pursuit of explaining and improving generalization in deep learning has elicited efforts both in regularization techniques as well as visualization techniques of the loss surface geometry. The latter is related to the intuition prevalent in the community that flatter local optima leads to lower generalization error…
Discretizes special surfaces using Koenigs nets.
This paper presents a new mathematical framework to analyze the loss functions of deep neural networks with ReLU functions. Furthermore, as as application of this theory, we prove that the loss functions can reconstruct the inputs of the training samples up to scalar multiplication (as vectors) and can provide the numb…
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…
One popular hypothesis of neural network generalization is that the flat local minima of loss surface in parameter space leads to good generalization. However, we demonstrate that loss surface in parameter space has no obvious relationship with generalization, especially under adversarial settings. Through visualizing …
Novel approach embeds loss tunnels in neural networks, revealing insights into their structure.
The work "Loss Landscape Sightseeing with Multi-Point Optimization" (Skorokhodov and Burtsev, 2019) demonstrated that one can empirically find arbitrary 2D binary patterns inside loss surfaces of popular neural networks. In this paper we prove that: (i) this is a general property of deep universal approximators; and (i…
Study finds Hilbert square of real surfaces can be maximal even when the surface has disconnected real locus.
We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…
The paper uses thermodynamics to improve machine learning representation quality.
Improves MRI-based brain surface reconstruction with minimal deformation energy loss.
We present multi-point optimization: an optimization technique that allows to train several models simultaneously without the need to keep the parameters of each one individually. The proposed method is used for a thorough empirical analysis of the loss landscape of neural networks. By extensive experiments on FashionM…
Building upon recent advances in entropy-regularized optimal transport, and upon Fenchel duality between measures and continuous functions , we propose a generalization of the logistic loss that incorporates a metric or cost between classes. Unlike previous attempts to use optimal transport distances for learning, our …
We establish an energy quantization result for sequences of Willmore surfaces when the underlying sequence of Riemann surfaces is degenerating in the moduli space. we notably exhibit a new residue which quantifies the potential loss of energy in collar regions. Thanks to these residues, we also prove compactness of Wil…
In the present work we study the behavior of sequences of smooth global isothermic immersions of a given closed surface and having a uniformly bounded total curvature. We prove that, if the conformal class of this sequence is bounded in the Moduli space of the surface, it weakly converges in W^{2,2} away from finitely …
It is widely conjectured that the reason that training algorithms for neural networks are successful because all local minima lead to similar performance, for example, see (LeCun et al., 2015, Choromanska et al., 2015, Dauphin et al., 2014). Performance is typically measured in terms of two metrics: training performanc…
Flat-minima optimizers improve neural network generalization.
Characterizes corridors in loss surfaces for gradient-based optimization.
Chinchilla Approach 2 biases neural scaling law estimates, leading to unnecessary compute costs.
Study optimizes CANN for actuarial tasks using RSM.
The early phase of training of deep neural networks is critical for their final performance. In this work, we study how the hyperparameters of stochastic gradient descent (SGD) used in the early phase of training affect the rest of the optimization trajectory. We argue for the existence of the "break-even" point on thi…
Current training methods for deep neural networks boil down to very high dimensional and non-convex optimization problems which are usually solved by a wide range of stochastic gradient descent methods. While these approaches tend to work in practice, there are still many gaps in the theoretical understanding of key as…
Much of the focus in machine learning research is placed in creating new architectures and optimization methods, but the overall loss function is seldom questioned. This paper interprets machine learning from a multi-objective optimization perspective, showing the limitations of the default linear combination of loss f…
New study shows deep networks generalize well due to loss surface geometry.
Transforms solutions of Davey-Stewartson II equation geometrically.
Gaussian process regression loses locality in high dimensions, affecting molecular energy surface fitting.
We extend the classical theory of isothermic surfaces in conformal 3-space, due to Bour, Christoffel, Darboux, Bianchi and others, to the more general context of submanifolds of symmetric -spaces with essentially no loss of integrable structure.
Graph networks struggle with multi-task learning due to varying property loss surface curvatures.
The Hessian of neural networks can be decomposed into a sum of two matrices: (i) the positive semidefinite generalized Gauss-Newton matrix G, and (ii) the matrix H containing negative eigenvalues. We observe that for wider networks, minimizing the loss with the gradient descent optimization maneuvers through surfaces o…
A neural flow method minimizes Willmore energy for 2-surfaces in 3D space.
We study the error landscape of deep linear and nonlinear neural networks with the squared error loss. Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete. For deep linear networks, we present necessary and s…
Study existence of harmonic and Dirac-harmonic maps from degenerating surfaces.
The loss functions of deep neural networks are complex and their geometric properties are not well understood. We show that the optima of these complex loss functions are in fact connected by simple curves over which training and test accuracy are nearly constant. We introduce a training procedure to discover these hig…
Focal loss reduces model curvature for better calibration.
Physics-informed neural network identifies and characterizes surface cracks in metals.
We introduce novel variants of momentum by incorporating the variance of the stochastic loss function. The variance characterizes the confidence or uncertainty of the local features of the averaged loss surface across the i.i.d. subsets of the training data defined by the mini-batches. We show two applications of the g…
In this paper we propose a method of obtaining points of extreme overfitting - parameters of modern neural networks, at which they demonstrate close to 100 % training accuracy, simultaneously with almost zero accuracy on the test sample. Despite the widespread opinion that the overwhelming majority of critical points o…