Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on how noise and variation-norm regularisation help shallow ReLU networks use fewer neurons.
This study uses continuous-time analysis to understand how momentum affects the optimisation of diagonal linear networks.
GD converges in unstable regimes, even with oscillatory behavior.
Bayesian posterior contraction rates improve with decreasing tails
We provide bounds on control learning error in stochastic systems.
The paper shows how label noise in training can lead to solutions that solve a Lasso program.
Gaussian processes (GPs) are Bayesian nonparametric generative models that provide interpretability of hyperparameters, admit closed-form expressions for training and inference, and are able to accurately represent uncertainty. To model general non-Gaussian data with complex correlation structure, GPs can be paired wit…
Study on GD and SGD over diagonal networks, focusing on stepsizes and regularisation.
Develops Bayesian filtering for online learning and related problems.
The paper analyzes fluctuations in ensemble models in high-dimensional settings.
A study on a surprising phase transition in model generalization error as parameters approach sample size.