Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

68137205273 · Jun 202019922001200920172026
48 results for zero-loss solutions

Zero loss is achievable in overparametrized DL networks under specific conditions.

problem Achieving zero loss in overparametrized deep learning networks.
method Determine sufficient conditions for zero loss attainability and present an explicit construction of zero loss minimizers.
result Explicit minimizers for zero loss in overparametrized DL networks are constructed without gradient descent.

Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.

problem Achieving zero loss minimizers in deep learning networks.
method Analysis of gradient descent algorithm in deep learning, focusing on underparametrized networks.
result Zero loss minimization cannot be achieved generically in deep learning networks.

One-pass SGD dynamics in overparameterized quadratic networks show slow escape from poor solutions.

problem Slow escape from poor generalization solutions in overparameterized neural networks.
method Analysis of one-pass SGD dynamics using ordinary differential equations for overlap matrices.
result Overparameterization only modestly accelerates escape from poor solutions.

Representing shapes as level sets of neural networks has been recently proved to be useful for different shape analysis and reconstruction tasks. So far, such representations were computed using either: (i) pre-computed implicit shape representations; or (ii) loss functions explicitly defined over the neural level sets…

2020-02-24abs ↗pdf ↗

Constructs classifiers for neural networks with specific data configurations.

problem Finding global minima of deep ReLU neural networks on sequentially separable data.
method Explicitly constructs zero loss neural network classifiers using cumulative parameters and truncation maps.
result Global minimizers can be described with a limited number of parameters based on the data structure.

New DP algorithms achieve near-optimal regret bounds for online learning problems.

problem Online learning problems with zero-loss solutions and differential privacy constraints.
method Developed new Differentially Private algorithms with near-optimal regret bounds.
result Achieved near-optimal regret bounds for various online prediction and convex optimization problems.

Generalization performance of classifiers in deep learning has recently become a subject of intense study. Deep models, typically over-parametrized, tend to fit the training data exactly. Despite this "overfitting", they perform well on test data, a phenomenon not yet fully understood. The first point of our paper is t…

2018-02-05abs ↗pdf ↗

We analyze multi-layer neural networks in the asymptotic regime of simultaneously (A) large network sizes and (B) large numbers of stochastic gradient descent training iterations. We rigorously establish the limiting behavior of the multi-layer neural network output. The limit procedure is valid for any number of hidde…

2019-03-11abs ↗pdf ↗

This study explains gradient flow dynamics in neural networks for small initialisation.

problem Understanding the training dynamics of neural networks for small initialisation.
method Analysis of gradient flow dynamics for one-hidden layer ReLU networks with orthogonal inputs.
result Gradient flow converges to zero loss and characterizes implicit bias towards minimum variation norm.

Wide neural networks converge linearly to zero loss with feature learning.

problem Optimizing wide neural networks with feature learning guarantees.
method Gradient flow analysis for wide shallow and multi-layer NNs.
result Training loss converges linearly to zero for wide NNs under GF, demonstrating feature learning and better generalization.

New method trains shallow neural networks with subquadratic width scaling.

problem Training shallow neural networks with optimal width scaling.
method Polyak-Lojasiewicz condition, smoothness, standard data assumptions, random matrix theory.
result Subquadratic scaling on network width with standard initialization strategies.

Augmenting a neural network with memory that can grow without growing the number of trained parameters is a recent powerful concept with many exciting applications. We propose a design of memory augmented neural networks (MANNs) called Labeled Memory Networks (LMNs) suited for tasks requiring online adaptation in class…

2017-07-05abs ↗pdf ↗

This paper explores how train-validation splits help in NAS to prevent overfitting.

problem NAS overfits with train-validation splits and needs better generalization guarantees.
method Established refined properties of validation loss and risk for NAS.
result NAS with train-validation splits can select the most generalizable model.

Newton's method converges faster than gradient descent in overparameterized neural networks.

problem Training neural networks efficiently in the overparameterized limit.
method Developed a convergence analysis for the regularized Newton method in this context.
result The NN training dynamics converge to the solution of a deterministic limit equation involving a Newton neural tangent kernel (NNTK).

Randomly trained neural networks can generalize well if there's a simpler underlying teacher model.

problem Why randomly trained neural networks generalize well despite interpolating training data.
method Examined a random neural network that interpolates training data and showed it generalizes well if there's a simpler underlying teacher model.
result Randomly trained neural networks can generalize well if there's a simpler underlying teacher model.

The paper proves ADL mechanisms face a trilemma and optimizes them for fairness, revenue, and exchange solvency.

problem The impossibility of a perpetual futures exchange achieving solvency, revenue, and fairness.
method Formal model of ADL, proving trilemma, and analyzing three ADL mechanisms.
result Optimized ADL mechanisms can reduce trader losses while maintaining exchange solvency.

Study shows how deep residual networks can be analyzed as shallow network ensembles for optimization.

problem Understanding why deep neural networks can be trained to zero loss despite non-convex optimization landscapes.
method Mean-field analysis of deep residual networks, focusing on their continuum limit as a two-layer network.
result Derives the first global convergence result for multilayer neural networks in the mean-field regime.

This paper investigates how large language models achieve neural collapse, a phenomenon linked to generalization.

problem Neural collapse in large language models under imbalanced and token-rich conditions.
method Empirical investigation of scaling and regularization effects on CLMs' progression towards neural collapse.
result Neural collapse properties develop with scale and regularization, linked to generalization in language modeling.

The paper analyzes the implicit bias of SGD near loss manifold and provides new insights.

problem Understanding the implicit bias of SGD near loss manifolds in overparametrized models.
method Adapting ideas from Katzenberger (1991) to analyze SGD dynamics using a stochastic differential equation (SDE).
result SGD with label noise locally decreases the sharpness of loss, leading to a global analysis of implicit bias.

The paper examines partial regularity of Lipschitz solutions to minimal surface system.

problem Understanding the regularity of solutions to the minimal surface system.
method Investigation of stationary, integral weak, and viscosity solutions; interior gradient estimate using maximum principle.
result Partial regularity results for Lipschitz solutions, including interior gradient estimate.

The paper constructs solutions to a critical Dirac equation on spheres.

problem Solving the critical Dirac equation on spheres with singularities.
method Constructing Delaunay-type solutions and another kind of singular solutions.
result The constructed solutions are building blocks for singular solutions on Spin manifolds.

We construct low regularity solutions of the vacuum Einstein constraint equations. In particular, on 3-manifolds we obtain solutions with metrics in $H^s\loc$ with s>32s>{3\over 2}. The theory of maximal asymptotically Euclidean solutions of the constraint equations descends completely the low regularity setting. Moreove…

2004-05-17abs ↗pdf ↗

Let n3n\ge 3 and m=n2n+2m=\frac{n-2}{n+2}. We construct 55-parameters, 44-parameters, 33-parameters ancient solutions of the equation vt=(vm)xx+vvmv_t=(v^m)_{xx}+v-v^m, v>0v>0, in R×(,T)\mathbb{R}\times (-\infty,T) for some TRT\in\mathbb{R}. This equation arises in the study of Yamabe flow. We obtain various properties of the ancient so…

2016-06-09abs ↗pdf ↗

We construct new ancient compact solutions to the Yamabe flow. Our solutions are rotationally symmetric and converge, as tt \to -\infty, to two self-similar complete non-compact solutions to the Yamabe flow moving in opposite directions. They are type I ancient solutions.

2015-09-29abs ↗pdf ↗

Paper shows how solutions to Allen-Cahn converge to multiphase mean curvature flow.

problem Convergence of Allen-Cahn solutions to multiphase mean curvature flow.
method Conditional convergence result of Allen-Cahn solutions to De Giorgi type BV-solutions of multiphase mean curvature flow.
result De Giorgi type BV-solutions are unique in a weak-strong sense.

We construct new ancient compact solutions to the Yamabe flow. Our solutions are rotationally symmetric and converge, as tt \to -\infty, to two self-similar complete non-compact solutions to the Yamabe flow moving in opposite directions. They are type I ancient solutions.

2016-01-20abs ↗pdf ↗

Generic level sets in mean curvature flow are BV solutions.

problem Understanding the behavior of level sets in mean curvature flow.
method Using the framework of sets of finite perimeter and distributional solutions, the paper extends Evans and Spruck's work.
result Generic level sets are distributional solutions with optimal energy dissipation rate.

Proves existence and uniqueness of viscosity solutions to complex Hessian equations on compact Hermitian manifolds.

problem Existence and uniqueness of viscosity solutions to complex Hessian equations.
method Proves existence and uniqueness using viscosity solutions and determinant domination conditions.
result Viscosity solutions exist and are unique under certain conditions.