Wide neural networks become linear, with constant tangent kernel, due to Hessian scaling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generalized linear models with nonlinear feature transformations are widely used for large-scale regression and classification problems with sparse inputs. Memorization of feature interactions through a wide set of cross-product feature transformations are effective and interpretable, while generalization requires more…
Wide neural networks converge linearly to zero loss with feature learning.
A longstanding goal in deep learning research has been to precisely characterize training and generalization. However, the often complex loss landscapes of neural networks have made a theory of learning dynamics elusive. In this work, we show that for wide neural networks the learning dynamics simplify considerably and…
Unified derivation of high-dimensional linear models using stochastic gradient descent.
This chapter introduces quaternion machine learning for 3D rotations.
The paper examines how deep linear neural networks behave as they become infinitely wide.
We prove that for an -layer fully-connected linear neural network, if the width of every hidden layer is , where and are the rank and the condition number of the input data, and is the output dimension, then gradient descent with Gaussi…
The paper analyzes knowledge distillation in wide neural networks, providing theoretical insights and practical implications.
Wide neural networks become linear, but adding bottlenecks makes them bilinear or multilinear.
Study of deep Stable neural networks with various activation functions.
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
LightOn OPUs accelerate randomized numerical linear algebra, reducing computational costs.
New method solves linear inverse problems using diffusion models.
Proposes using MLP for predicting optimal penalty in changepoint detection.
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer, or 2) at least as wide as the output layer. This result is the strongest possibl…
Wide neural networks with weight decay exhibit neural collapse.
Study compares dropout and l2 regularization in linear models.
2 Diabetes is a leading worldwide public health concern, and its increasing prevalence has significant health and economic importance in all nations. The condition is a multifactorial disorder with a complex aetiology. The genetic determinants remain largely elusive, with only a handful of identified candidate genes. G…
In machine learning and data mining, linear models have been widely used to model the response as parametric linear functions of the predictors. To relax such stringent assumptions made by parametric linear models, additive models consider the response to be a summation of unknown transformations applied on the predict…
Given a linear regression setting, Iterative Least Trimmed Squares (ILTS) involves alternating between (a) selecting the subset of samples with lowest current loss, and (b) re-fitting the linear model only on that subset. Both steps are very fast and simple. In this paper we analyze ILTS in the setting of mixed linear …
New lower bounds for linear classification problems in high dimensions.
LIFE framework improves model accuracy and interpretability.
Region-specific linear models are widely used in practical applications because of their non-linear but highly interpretable model representations. One of the key challenges in their use is non-convexity in simultaneous optimization of regions and region-specific models. This paper proposes novel convex region-specific…
Estimates GLMs robustly against label corruptions.
Wide networks with polynomial activations have proven asymptotic behavior.
Use simplified layerwise linear models to understand neural dynamics.
We find that the CAPM fails to explain the small firm effect even if its non-parametric form is used which allows time-varying risk and non-linearity in the pricing function. Furthermore, the linearity of the CAPM can be rejected, thus the widely used risk and performance measures, the beta and the alpha, are biased an…
GCNIII combines Wide & Deep for better node classification.
Adequate evaluation of an information retrieval system to estimate future performance is a crucial task. Area under the ROC curve (AUC) is widely used to evaluate the generalization of a retrieval system. However, the objective function optimized in many retrieval systems is the error rate and not the AUC value. This p…
Paper presents a machine learning method to improve significance tests for misspecified linear models.
Gaussian belief propagation (BP) has been widely used for distributed inference in large-scale networks such as the smart grid, sensor networks, and social networks, where local measurements/observations are scattered over a wide geographical area. One particular case is when two neighboring agents share a common obser…
Efficiently solves inverse PDE problems with Gaussian processes.
We study online linear regression problems in a distributed setting, where the data is spread over a network. In each round, each network node proposes a linear predictor, with the objective of fitting the \emph{network-wide} data. It then updates its predictor for the next round according to the received local feedbac…
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training -hidden-layer linear residual networks (ResNets). We prove that for training deep residual networks with certain linear transformations at input and output layers, which are fixed throughout training, both GD and SGD…
Random non-linear Fourier features have recently shown remarkable performance in a wide-range of regression and classification applications. Motivated by this success, this article focuses on a sparse non-linear Fourier feature (NFF) model. We provide a characterization of the sufficient number of data points that guar…
We learn linear models from nonlinear systems using multiple trajectories and regularization.
A linear multi-factor model is one of the most important tools in equity portfolio management. The linear multi-factor models are widely used because they can be easily interpreted. However, financial markets are not linear and their accuracy is limited. Recently, deep learning methods were proposed to predict stock re…
This study uses neural networks to solve interpolation problems with sparse, infinitely wide layers.
Cyclic coordinate descent identifies models in finite time and converges linearly.
Paper proposes a new method for sparse spectral clustering on Stiefel manifold.
New adaptive models improve prediction accuracy with missing data.
Selective neural network improves credit risk prediction while maintaining interpretability.
Efficient algorithms recover two sparse models from a mix of linear queries.
Sub-sampling is a common and often effective method to deal with the computational challenges of large datasets. However, for most statistical models, there is no well-motivated approach for drawing a non-uniform subsample. We show that the concept of an asymptotically linear estimator and the associated influence func…
The autoencoder is an effective unsupervised learning model which is widely used in deep learning. It is well known that an autoencoder with a single fully-connected hidden layer, a linear activation function and a squared error cost function trains weights that span the same subspace as the one spanned by the principa…
Gaussian processes (GP) are a widely used model for regression problems in supervised machine learning. Implementation of GP regression typically requires logic gates. We show that the quantum linear systems algorithm [Harrow et al., Phys. Rev. Lett. 103, 150502 (2009)] can be applied to Gaussian process regre…
We identify linear models from nonlinear systems with initialization constraints.