Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

86172258344 · Jun 202019922001200920172026
48 results for Loss Surface

We propose to study neural networks' loss surfaces by methods of topological data analysis. We suggest to apply barcodes of Morse complexes to explore topology of loss surfaces. An algorithm for calculations of the loss function's barcodes of local minima is described. We have conducted experiments for calculating barc…

2019-11-29abs ↗pdf ↗

Paper discovers simplicial complexes connecting trained models for improved ensembling.

problem Improving robustness and accuracy of deep learning ensembles.
method Identifies mode-connecting simplicial complexes on loss surfaces.
result Efficiently builds simplicial complexes for ensembling, outperforming independent ensembles.

New tools reveal simple structure in complex hyperparameter loss surfaces near optima.

problem Understanding the behavior of hyperparameter loss surfaces near model optima.
method Developed a novel technique based on random search to uncover asymptotic features of the loss surface.
result Random search within the asymptotic regime yields a new distribution with parameters defining the loss surface.

Piecewise linear activations create many spurious local minima in neural networks.

problem Understanding the loss surface of neural networks with piecewise linear activations.
method Proved the existence of infinite spurious local minima and partitioned the loss surface into smooth cells.
result Piecewise linear activations create many spurious local minima that are invariant under a continuous path.

Study compact Willmore surfaces without complex structure convergence, computing energy loss and geodesic lengths.

problem Compactness of Willmore surfaces without complex structure convergence.
method Compute energy loss in neck and geodesic lengths in Grassmannian G(2,n)G(2,n).
result Limit of Gauss map image is a geodesic in G(2,n)G(2,n) with computable length.

It has been argued in the past that high-dimensional neural networks do not exhibit local minima capable of trapping an optimisation algorithm. However, the relationship between loss surface modality and the neural architecture parameters, such as the number of hidden neurons per layer and the number of hidden layers, …

2019-05-24abs ↗pdf ↗

The pursuit of explaining and improving generalization in deep learning has elicited efforts both in regularization techniques as well as visualization techniques of the loss surface geometry. The latter is related to the intuition prevalent in the community that flatter local optima leads to lower generalization error…

2019-07-22abs ↗pdf ↗

This paper presents a new mathematical framework to analyze the loss functions of deep neural networks with ReLU functions. Furthermore, as as application of this theory, we prove that the loss functions can reconstruct the inputs of the training samples up to scalar multiplication (as vectors) and can provide the numb…

2018-05-18abs ↗pdf ↗

In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…

2017-02-27abs ↗pdf ↗

Novel approach embeds loss tunnels in neural networks, revealing insights into their structure.

problem Understanding the structure of neural network loss surfaces, especially low-loss tunnels.
method Directly embedding loss tunnels into the loss landscape of neural networks.
result Improved insights into the length and structure of loss tunnels, and better subspace inference in Bayesian neural networks.

Study finds Hilbert square of real surfaces can be maximal even when the surface has disconnected real locus.

problem Exploring conditions for maximality of Hilbert square of real surfaces.
method Analyzing Hilbert square of maximal real surfaces and examining specific examples.
result Hilbert square can be maximal even for surfaces with disconnected real locus.

We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…

2018-02-24abs ↗pdf ↗

The paper uses thermodynamics to improve machine learning representation quality.

problem Improving the quality of learned representations for transfer learning.
method Formal connection with thermodynamics, iso-classification process, traversing the equilibrium surface.
result Demonstrates how to transfer representations while keeping classification loss constant.

Improves MRI-based brain surface reconstruction with minimal deformation energy loss.

problem Ensuring optimal deformation energy and consistency in learning-based cortical surface reconstruction.
method Design and implementation of a Minimal Energy Deformation (MED) loss in the V2C-Flow model.
result Significant improvements in training consistency and reproducibility without sacrificing reconstruction accuracy and topological correctness.

We present multi-point optimization: an optimization technique that allows to train several models simultaneously without the need to keep the parameters of each one individually. The proposed method is used for a thorough empirical analysis of the loss landscape of neural networks. By extensive experiments on FashionM…

2019-10-09abs ↗pdf ↗

Building upon recent advances in entropy-regularized optimal transport, and upon Fenchel duality between measures and continuous functions , we propose a generalization of the logistic loss that incorporates a metric or cost between classes. Unlike previous attempts to use optimal transport distances for learning, our …

2019-05-15abs ↗pdf ↗

In the present work we study the behavior of sequences of smooth global isothermic immersions of a given closed surface and having a uniformly bounded total curvature. We prove that, if the conformal class of this sequence is bounded in the Moduli space of the surface, it weakly converges in W^{2,2} away from finitely …

2012-02-06abs ↗pdf ↗

Characterizes corridors in loss surfaces for gradient-based optimization.

problem Understanding and mitigating training instabilities in gradient-based optimization.
method Characterizes corridors as regions where gradient descent and gradient flow trajectories are linearly related.
result Corridors indicate regions without implicit regularization effects, leading to better learning rate adaptation schemes.

Chinchilla Approach 2 biases neural scaling law estimates, leading to unnecessary compute costs.

problem Systematic biases in Chinchilla Approach 2's parabolic fits of neural scaling laws.
method Analyzes three sources of error: IsoFLOP sampling grid width, uncentered sampling, and loss surface asymmetry.
result Chinchilla Approach 3 largely eliminates these biases, offering a more convenient or scalable alternative.

Gaussian process regression loses locality in high dimensions, affecting molecular energy surface fitting.

problem Loss of locality in high-dimensional Gaussian process regression.
method Analysis of Matern family kernels and multi-zeta basis functions.
result The property of locality disappears in high dimensions, impacting regression quality.

We extend the classical theory of isothermic surfaces in conformal 3-space, due to Bour, Christoffel, Darboux, Bianchi and others, to the more general context of submanifolds of symmetric RR-spaces with essentially no loss of integrable structure.

2009-06-09abs ↗pdf ↗

Graph networks struggle with multi-task learning due to varying property loss surface curvatures.

problem Graph networks underperform in multi-task learning for crystal and molecule properties.
method Assessed curvature of property loss surfaces via spectral properties of Hessians, matrix-free using randomized numerical linear algebra.
result Varying curvature of property loss surfaces explains graph networks' multi-task learning inefficiency.

A neural flow method minimizes Willmore energy for 2-surfaces in 3D space.

problem Minimizing Willmore energy for closed oriented 2-surfaces in 3D space.
method Introducing neural Willmore flow to model and minimize the Willmore energy using neural architectures.
result The neural flow reproduces expected round sphere and Clifford torus for genus 0 and 1 surfaces, respectively, and finds minimal Willmore surfaces for genus 2.

We study the error landscape of deep linear and nonlinear neural networks with the squared error loss. Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete. For deep linear networks, we present necessary and s…

2017-07-08abs ↗pdf ↗

Study existence of harmonic and Dirac-harmonic maps from degenerating surfaces.

problem Existence of harmonic and Dirac-harmonic maps from degenerating surfaces.
method Using the Sacks and Uhlenbeck scheme, analyze a sequence of maps from degenerating surfaces to non-positive curved manifolds.
result Existence of limiting harmonic and Dirac-harmonic maps under certain conditions.

The loss functions of deep neural networks are complex and their geometric properties are not well understood. We show that the optima of these complex loss functions are in fact connected by simple curves over which training and test accuracy are nearly constant. We introduce a training procedure to discover these hig…

2018-02-27abs ↗pdf ↗

Physics-informed neural network identifies and characterizes surface cracks in metals.

problem Identifying and characterizing surface-breaking cracks in metals using ultrasound.
method Physics-informed neural network (PINN) trained with ultrasonic surface wave data and adaptive activation functions.
result PINN accurately estimates the speed of sound and identifies crack locations in metals.

We introduce novel variants of momentum by incorporating the variance of the stochastic loss function. The variance characterizes the confidence or uncertainty of the local features of the averaged loss surface across the i.i.d. subsets of the training data defined by the mini-batches. We show two applications of the g…

2019-05-30abs ↗pdf ↗

In this paper we propose a method of obtaining points of extreme overfitting - parameters of modern neural networks, at which they demonstrate close to 100 % training accuracy, simultaneously with almost zero accuracy on the test sample. Despite the widespread opinion that the overwhelming majority of critical points o…

2019-06-14abs ↗pdf ↗