SGD vs quasi-Newton optimization in neural networks: different landscapes, different generalizability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Persistence landscapes map persistence diagrams into a function space, which may often be taken to be a Banach space or even a Hilbert space. In the latter case, it is a feature map and there is an associated kernel. The main advantage of this summary is that it allows one to apply tools from statistics and machine lea…
Deep ReLU networks with extra parameters have mostly good loss landscapes.
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce mi…
Training an artificial neural network involves an optimization process over the landscape defined by the cost (loss) as a function of the network parameters. We explore these landscapes using optimisation tools developed for potential energy landscapes in molecular science. The number of local minima and transition sta…
We explore the energy landscape of a simple neural network. In particular, we expand upon previous work demonstrating that the empirical complexity of fitted neural networks is vastly less than a naive parameter count would suggest and that this implicit regularization is actually beneficial for generalization from fit…
It is widely observed that deep learning models with learned parameters generalize well, even with much more model parameters than the number of training samples. We systematically investigate the underlying reasons why deep neural networks often generalize well, and reveal the difference between the minima (with the s…
The complex and computationally expensive nature of landscape evolution models pose significant challenges in the inference and optimisation of unknown parameters. Bayesian inference provides a methodology for estimation and uncertainty quantification of unknown model parameters. In our previous work, we developed para…
Wide deep neural networks are easy to optimize without constraints.
Studying SGD on deep neural networks using diffusion maps.
New framework to understand and exploit curvature in deep learning loss landscapes.
The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the core technique from these papers can be used to remove all bad local minima from any loss landscape, so long as the global minimum has a loss o…
We investigate the structure of the profit landscape obtained from the most basic, fluctuation based, trading strategy applied for the daily stock price data. The strategy is parameterized by only two variables, p and q. Stocks are sold and bought if the log return is bigger than p and less than -q, respectively. Repet…
Almost all local minima in neural networks are strongly convex.
Neural networks' optimization dynamics are confined to a single basin despite connected basins in the loss landscape.
New method simplifies optimization landscapes by transforming saddle points.
Study finds non-IID data causes FL performance issues.
Improved Langevin Monte Carlo reduces energy barriers for faster optimization.
Overparametrization improves QNN trainability by reducing spurious local minima.
Many recently trained neural networks employ large numbers of parameters to achieve good performance. One may intuitively use the number of parameters required as a rough gauge of the difficulty of a problem. But how accurate are such notions? How many parameters are really needed? In this paper we attempt to answer th…
Monotonic Linear Interpolation property in neural networks persists despite non-convexity.
Most high-dimensional estimation and prediction methods propose to minimize a cost function (empirical risk) that is written as a sum of losses associated to each data point. In this paper we focus on the case of non-convex losses, which is practically important but still poorly understood. Classical empirical process …
We present multi-point optimization: an optimization technique that allows to train several models simultaneously without the need to keep the parameters of each one individually. The proposed method is used for a thorough empirical analysis of the loss landscape of neural networks. By extensive experiments on FashionM…
In many high-dimensional estimation problems the main task consists in minimizing a cost function, which is often strongly non-convex when scanned in the space of parameters to be estimated. A standard solution to flatten the corresponding rough landscape consists in summing the losses associated to different data poin…
Gradient descent variants improve phase retrieval accuracy.
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
Data-driven model shows deep learning weights behave like a liquid.
We explore some mathematical features of the loss landscape of overparameterized neural networks. A priori one might imagine that the loss function looks like a typical function from to - in particular, nonconvex, with discrete global minima. In this paper, we prove that in at least one impo…
We apply a simple trading strategy for various time series of real and artificial stock prices to understand the origin of fractality observed in the resulting profit landscapes. The strategy contains only two parameters and , and the sell (buy) decision is made when the log return is larger (smaller) than (…
Paper analyzes Transformer learning dynamics, proving benign landscape for in-context learning.
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an -layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rat…
This paper presents approximate confidence intervals for each function of parameters in a Banach space based on a bootstrap algorithm. We apply kernel density approach to estimate the persistence landscape. In addition, we evaluate the quality distribution function estimator of random variables using integrated mean sq…
Three training regimes found for scale-invariant neural networks on the sphere.
Variational quantum computing faces a flat optimization landscape problem.
The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence. However, the loss functions of deep neural networks (especially nonlinear networks) are still far from being well understood from a theoretical aspect. In this pa…
Deep neural networks are workhorse models in machine learning with multiple layers of non-linear functions composed in series. Their loss function is highly non-convex, yet empirically even gradient descent minimisation is sufficient to arrive at accurate and predictive models. It is hitherto unknown why are deep neura…
Reviews recent findings on neural network landscapes.
Large learning rates cause parameter instability, leading to better generalization.
In this work we analyse quantitatively the interplay between the loss landscape and performance of descent algorithms in a prototypical inference problem, the spiked matrix-tensor model. We study a loss function that is the negative log-likelihood of the model. We analyse the number of local minima at a fixed distance …
We analyze the landscape of empirical risk minimization for high-dimensional models, predicting phase transitions and critical point properties.
This work proposes a mathematical framework for loss landscapes and optimization in deep neural networks.
Efficient MCMC sampling in Bayesian neural networks by exploiting symmetries.
The paper investigates what enables successful transfer learning and separates feature reuse from data statistics.
We propose a general theory for studying the \xl{landscape} of nonconvex \xl{optimization} with underlying symmetric structures \tz{for a class of machine learning problems (e.g., low-rank matrix factorization, phase retrieval, and deep linear neural networks)}. In specific, we characterize the locations of stationary …
The study examines when MAML's objective has a benign landscape.
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
Black holes offer insights into machine learning's loss landscapes.