Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper constructs minimizers for deep learning networks and analyzes their geometric structure.
The paper constructs upper bounds for cost minimization in shallow neural networks.
Zero loss is achievable in overparametrized DL networks under specific conditions.
COCOA can learn new tasks without forgetting previously learned ones in a distributed setting.
New theory explains how overparametrized neural networks generalize well without bias-variance trade-off.
Overparametrization improves QNN trainability by reducing spurious local minima.
Noise can affect the overparametrization of QNNs, enabling new directions but also suppressing sensitivity.
Adversarial training improves linear regression solutions, offering robustness against small perturbations.
Gradient descent with geometrically adapted metrics drives cost to global minimum at uniform rate.
Bagging stabilizes linear interpolators, improving their generalization performance.