Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
WSD schedule improves model training efficiency by adapting learning rates dynamically.
New framework reveals thermodynamic principles for LLM training.
Quantization-aware training can recover accuracy lost by post-training quantization.
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that from any point in parameter space there exists a continuous path on which the cross-entropy loss is non-increasing and gets arbitrarily clos…
Tilting loss functions improves machine learning performance.
Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenval…
We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
Study identifies new stable climate states in climate model.
Study the landscape of Lipschitz functions between manifolds using persistent homology.
Large SGD step sizes lead to sparse feature learning in neural networks.
Paper presents a GAN model for realistic river image synthesis.
Artificial Neural Network (ANN) based model is a computational approach commonly used for modeling the complex relationships between input and output parameters. Prediction of the flow rate of a river is a requisite for any successful water resource management and river basin planning. In the current survey, the effect…
Combining insights from machine learning and quantum Monte Carlo, the stochastic reconfiguration method with neural network Ansatz states is a promising new direction for high-precision ground state estimation of quantum many-body problems. Even though this method works well in practice, little is known about the learn…
We solve the optimization of two-layer ReLU networks using convex math.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
HydroNets use river structure to improve hydrologic predictions.
New method SF-AdamW trains large models without decay phases or memory overhead.
We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while keeping other modules fixed at the trained parameters, results in a large drop in the network's performance. Our analysis reveals interestin…
Online algorithms for identifying river pollution sources.
Deep learning improves probabilistic river discharge forecasting for hydroelectric power.
We consider remodeling the planar search patterns, in the presence of the river-type perturbation represented by the weak vector field, basing on the time-optimal paths as Finslerian solutions to the Zermelo navigation problem via Randers metric.
Study uses AUVs and RL to map river plumes over multiple days.
Modeling daily river flow distribution with seasonal and long-term trends.
Stochastic gradient descent (SGD) forms the core optimization method for deep neural networks. While some theoretical progress has been made, it still remains unclear why SGD leads the learning dynamics in overparameterized networks to solutions that generalize well. Here we show that for overparameterized networks wit…
New algorithm ATENT improves adversarial robustness in neural networks.
Learning to optimize - the idea that we can learn from data algorithms that optimize a numerical criterion - has recently been at the heart of a growing number of research efforts. One of the most challenging issues within this approach is to learn a policy that is able to optimize over classes of functions that are fa…
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large…
Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD). Recent work conjectures that SGD with proper configuration is able to find wide and flat local minima, which have been proposed to be associated with good gener…
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
A new pruning method reduces neural network computation without retraining.
A D-Wave quantum annealer (QA) having a 2048 qubit lattice, with no missing qubits and couplings, allowed embedding of a complete graph of a Restricted Boltzmann Machine (RBM). A handwritten digit OptDigits data set having 8x7 pixels of visible units was used to train the RBM using a classical Contrastive Divergence. E…
A new Kolmogorov-Arnold network improves function approximation and optimization.
The study analyzes river water quality using statistical and machine learning methods.
A common difficulty in applications of machine learning is the lack of any general principle for guiding the choices of key parameters of the underlying neural network. Focusing on a class of recurrent neural networks - reservoir computing systems that have recently been exploited for model-free prediction of nonlinear…
This paper examines SVB's failure and its impact on bank stocks.
Delivering useful hydrological forecasts is critical for urban and agricultural water management, hydropower generation, flood protection and management, drought mitigation and alleviation, and river basin planning and management, among others. In this work, we present and appraise a new simple and flexible methodology…
Flooding is a destructive and dangerous hazard and climate change appears to be increasing the frequency of catastrophic flooding events around the world. Physics-based flood models are costly to calibrate and are rarely generalizable across different river basins, as model outputs are sensitive to site-specific parame…
Learning hydrologic models for accurate riverine flood prediction at scale is a challenge of great importance. One of the key difficulties is the need to rely on in-situ river discharge measurements, which can be quite scarce and unreliable, particularly in regions where floods cause the most damage every year. Accordi…
Proposes a method to partition univariate data into unimodal subsets.
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima. In a network of hidden layers with neurons in layers $k = 1, \ldots, …
Model shows how capital accumulation can lead to poverty traps and well-being states.
Method reduces model bias in water temperature prediction using physics-guided GNNs.
CausalRivers benchmarks causal discovery methods on real-world river discharge data.
Dataset for rainfall modeling in central Europe from 1981-2011.