Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

0111 · Feb 202019922001200920172026
11 results for overparameterisation

Overparameterisation can limit the benefits of curriculum learning in neural networks.

problem The ineffectiveness of curriculum learning in deep learning applications.
method Analytical study connecting curriculum learning and overparameterisation in an online learning setting for a 2-layer network in the XOR-like Gaussian Mixture problem.
result High degree of overparameterisation can limit the benefit from curricula.

Depth uncertainty networks don't improve with bias correction, contrary to expectations.

problem Improving performance in active learning with overparameterised models like NNs.
method Depth uncertainty networks, compared to underparameterised models, show no improvement in performance with bias correction.
result Depth uncertainty networks do not improve with bias correction, unlike underparameterised models.

New stability bounds for GD in overparameterised shallow nets without NTK assumptions.

problem Generalisation and excess risk bounds for shallow neural networks.
method Oracle inequalities and stability analysis of GD without kernelisation.
result Oracle type bounds reveal GD's generalisation is controlled by an interpolating network with shortest GD path.

Study reveals sharp characterisation of local minima in neural network loss landscapes.

problem Characterizing local minima in high-dimensional two-layer ReLU neural networks.
method Exact low-dimensional representation of local minima using summary statistics and link with one-pass SGD dynamics.
result Local minima in overparameterized neural networks form discrete families with varying stability and reachability.

SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.

problem Understanding why overparameterized deep nets generalize well despite fitting training data perfectly.
method Empirical investigation of PSGD(fS)P_{SGD}(f\mid S) and PB(fS)P_B(f\mid S) for various architectures and datasets.
result The probability of SGD-converging on a function consistent with training data correlates well with the Bayesian posterior probability of expressing that function.

New research shows non-bottlenecked autoencoders can outperform bottlenecked ones for anomaly detection.

problem The necessity of a bottleneck in autoencoders for anomaly detection.
method Investigated two ways to remove bottlenecks: overparameterising the latent layer and introducing skip connections. Carried out extensive experiments on various AE types and datasets.
result Non-bottlenecked autoencoders can outperform bottlenecked ones, improving anomaly detection performance.

New bounds link flat minima to good generalisation in overparameterized models.

problem Understanding the relationship between flat minima and generalisation in overparameterized machine learning models.
method Combining PAC-Bayes, Poincaré, and Log-Sobolev inequalities to derive generalisation bounds involving gradient terms.
result Flat minima positively influence generalisation performance, highlighting the benefits of the optimisation phase.

Early alignment in neural networks leads to sparse representations but hinders convergence.

problem The implicit bias of gradient descent during early training phases.
method Quantitative description of early alignment phase in small initialisation, one hidden layer networks.
result Early alignment induces a sparse representation but also hinders convergence to global minima.

Analyzes SGD dynamics in two-layer networks, bridging different regimes.

problem Understanding SGD dynamics in high-dimensional and mean-field settings.
method Rigorous analysis via deterministic low-dimensional description of sufficient statistics.
result Infinite-width dynamics remains close to a low-dimensional subspace.

A new geometric concept, the dead direction, bridges singular learning theory and information geometry.

problem The gap between singular learning theory and information geometry.
method Introducing the dead direction, a unit vector along degenerating Fisher metric, and showing its KL order can be recovered.
result The KL order of the dead direction can be recovered as the decay rate of the directional Fisher curvature, providing a handle on singular geometry.