Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…
Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large…
Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD). Recent work conjectures that SGD with proper configuration is able to find wide and flat local minima, which have been proposed to be associated with good gener…
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
A new pruning method reduces neural network computation without retraining.
A D-Wave quantum annealer (QA) having a 2048 qubit lattice, with no missing qubits and couplings, allowed embedding of a complete graph of a Restricted Boltzmann Machine (RBM). A handwritten digit OptDigits data set having 8x7 pixels of visible units was used to train the RBM using a classical Contrastive Divergence. E…
A new Kolmogorov-Arnold network improves function approximation and optimization.
A common difficulty in applications of machine learning is the lack of any general principle for guiding the choices of key parameters of the underlying neural network. Focusing on a class of recurrent neural networks - reservoir computing systems that have recently been exploited for model-free prediction of nonlinear…
We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that from any point in parameter space there exists a continuous path on which the cross-entropy loss is non-increasing and gets arbitrarily clos…
This paper examines SVB's failure and its impact on bank stocks.
Proposes a method to partition univariate data into unimodal subsets.
New framework reveals thermodynamic principles for LLM training.
WSD schedule improves model training efficiency by adapting learning rates dynamically.
HyPV-LEAD detects cryptocurrency anomalies proactively, improving financial security.
Quantization-aware training can recover accuracy lost by post-training quantization.
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenval…
Solves local minima problems on smooth manifolds.
Unbalanced data arises in many learning tasks such as clustering of multi-class data, hierarchical divisive clustering and semisupervised learning. Graph-based approaches are popular tools for these problems. Graph construction is an important aspect of graph-based learning. We show that graph-based algorithms can fail…
Designs chiral photonic structures using machine learning for efficient optical properties.
Tilting loss functions improves machine learning performance.
Stablecoin liquidity was affected by the SVB collapse, with USDC's transparency leading to market reactions.
Large SGD step sizes lead to sparse feature learning in neural networks.
A new pruning method finds sparse minimizers in flat regions of deep neural networks.
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
Cryptos remained resilient after SVB's collapse, contrary to expectations.
Piecewise linear activations create many spurious local minima in neural networks.
Model shows how banks' hidden-to-maturity accounting can mask run risk and lead to financial instability.
D-Wave quantum annealing fails to improve sampling quality from RBMs compared to Gibbs sampling.
SALR improves deep learning generalization by dynamically adjusting learning rates.
We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while keeping other modules fixed at the trained parameters, results in a large drop in the network's performance. Our analysis reveals interestin…
Study the landscape of Lipschitz functions between manifolds using persistent homology.
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima. In a network of hidden layers with neurons in layers $k = 1, \ldots, …
Stochastic gradient descent (SGD) forms the core optimization method for deep neural networks. While some theoretical progress has been made, it still remains unclear why SGD leads the learning dynamics in overparameterized networks to solutions that generalize well. Here we show that for overparameterized networks wit…
Model shows how capital accumulation can lead to poverty traps and well-being states.
New algorithm ATENT improves adversarial robustness in neural networks.
We build a model using Gaussian processes to infer a spatio-temporal vector field from observed agent trajectories. Significant landmarks or influence points in agent surroundings are jointly derived through vector calculus operations that indicate presence of sources and sinks. We evaluate these influence points by us…
A new AI optimization method uses energy-conserving dynamics inspired by Born-Infeld theory.
Randomized algorithms that base iteration-level decisions on samples from some pool are ubiquitous in machine learning and optimization. Examples include stochastic gradient descent and randomized coordinate descent. This paper makes progress at theoretically evaluating the difference in performance between sampling wi…
This paper presents an algorithm for a complete and efficient calibration of the Heston stochastic volatility model. We express the calibration as a nonlinear least squares problem. We exploit a suitable representation of the Heston characteristic function and modify it to avoid discontinuities caused by branch switchi…
DiMS sampler explores neural network loss minima via dissipative dynamics.
Combining insights from machine learning and quantum Monte Carlo, the stochastic reconfiguration method with neural network Ansatz states is a promising new direction for high-precision ground state estimation of quantum many-body problems. Even though this method works well in practice, little is known about the learn…
Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of the data, namely a human-readable chart of the probability density from which t…
Groups of firms often achieve a competitive advantage through the formation of geo-industrial clusters. Although many exemplary clusters, such as Hollywood or Silicon Valley, have been frequently studied, systematic approaches to identify and analyze the hierarchical structure of the geo-industrial clusters at the glob…
Study identifies new stable climate states in climate model.
Bayesian model averaging (BMA) is the state of the art approach for overcoming model uncertainty. Yet, especially on small data sets, the results yielded by BMA might be sensitive to the prior over the models. Credal Model Averaging (CMA) addresses this problem by substituting the single prior over the models by a set …