Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3.3%6.7%10.0%13.3% · May 199619922001200920172026
48 results for small initialization

The paper studies neural networks' convergence near origin and saddle points.

problem Directional convergence of neural networks near small initializations and saddle points.
method Gradient flow dynamics analysis of two-homogeneous neural networks.
result Neural networks' weights approximately converge in direction to KKT points for small initializations.

We show the existence of a global unique and analytic solution for the mean curvature flow, the surface diffusion flow and the Willmore flow of entire graphs for Lipschitz initial data with small Lipschitz norm. We also show the existence of a global unique and analytic solution to the Ricci-DeTurck flow on euclidean s…

2009-02-09abs ↗pdf ↗

Early training of deep neural networks leads to small, directionally converging weights.

problem Training dynamics of deep homogeneous neural networks with small initializations.
method Gradient flow analysis and study of KKT points for neural correlation function.
result Weights converge in direction to KKT points during early training stages.

We show that on Kahler manifolds M with c_1(M)=0 the Calabi flow converges to a constant scalar curvature metric if the initial Calabi energy is sufficiently small. We prove a similar result on manifolds with c_1(M)<0 if the Kahler class is close to the canonical class.

2006-08-07abs ↗pdf ↗

Gradient descent with small random init mimics spectral methods for low-rank matrix recovery.

problem Reconstructing a low-rank matrix from few measurements.
method Gradient descent with small random initialization followed by a few iterations.
result Gradient descent from small random init converges to a well-generalizing solution.

Proves global existence and uniqueness of solutions for Einstein-scalar-field equations.

problem Global existence and uniqueness of solutions for specific Einstein-scalar-field equations.
method Proves global existence and uniqueness of classical solutions with small initial data and wake-like decaying null infinity.
result Global existence and uniqueness of solutions for the equations with wake-like decaying null infinity.

We define and study the harmonic heat flow for almost complex structures which are compatible with a Riemannian structure (M,g)(M, g). This is a tensor-valued version of harmonic map heat flow. We prove that if the initial almost complex structure JJ has small energy (depending on the norm J|\nabla J|), then the flow ex…

2019-07-29abs ↗pdf ↗

Gradient descent with small initialization solves matrix completion without regularization.

problem Symmetric matrix completion from observed entries.
method Vanilla gradient descent with small initialization.
result GD converges to the ground truth matrix without regularization in over-parameterized scenario.

Deep linear networks minimize sharpness, avoiding large eigenvalues.

problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.

Analyzes Willmore flow for graphs with boundary data, proving existence and convergence.

problem Willmore flow of graphs with boundary conditions over bounded domains.
method Developed low-regularity theory, reformulated graphical equation, used time-weighted parabolic Hölder spaces.
result Proved short-time and global existence for initial data in C1+α(Ω)C^{1+α}(\overlineΩ) and Lipschitz, with exponential convergence.

Generative model initializes 2-layer network weights for small datasets.

problem Approximating functions with 2-layer networks using small datasets and gradient-based training.
method Initialize hidden weights with a learned proposal distribution parameterized as a deep generative model. Refine with gradient-based post-processing and regularization.
result Demonstrates effectiveness of the approach with numerical examples.

Standard practice in training neural networks involves initializing the weights in an independent fashion. The results of recent work suggest that feature "diversity" at initialization plays an important role in training the network. However, other initialization schemes with reduced feature diversity have also been sh…

2019-12-11abs ↗pdf ↗

We establish an optimal gluing construction for general relativistic initial data sets. The construction is optimal in two distinct ways. First, it applies to generic initial data sets and the required (generically satisfied) hypotheses are geometrically and physically natural. Secondly, the construction is completely …

2004-09-10abs ↗pdf ↗

We study the problem of inviscid slightly compressible fluids in a bounded domain. We find a unique solution to the initial-boundary value problem and show that it is near the analogous solution for an incompressible fluid provided the initial conditions for the two problems are close. In particular, the divergence of …

2013-09-02abs ↗pdf ↗

Unique solutions found for wave-like decaying null infinity equations.

problem Wave-like decaying null infinity equations with spherically symmetric Einstein-scalar-field.
method Local and global unique solutions for small initial data.
result Sharp decaying condition for unique solutions.

This paper presents a phase diagram for two-layer neural networks under different initialization scales.

problem Understanding the behavior of neural networks under varying scales of initialization.
method Analysis of a phase diagram for two-layer neural networks.
result Condensation of weight vectors on isolated orientations during training.

New initialization schemes preserve fractional moments of weights in deep networks, improving training and test performance.

problem Heavy-tailed distribution of stochastic gradients in DNNs during training.
method Developed initialization schemes that preserve any given fractional moment of order s < 2 over layers for various activations.
result The network output admits a heavy-tailed distribution with finite moments, improving training and test performance.

New method for better initial centers in clustering with improved accuracy and privacy.

problem Improving the quality of clustering centers in metric spaces.
method HST initialization based on metric embedding tree structure, combined with efficient search algorithm and DP extension.
result HST initialization produces better initial centers than kk-median++ with comparable efficiency and improved privacy.

Paper shows robustness of gradient descent in matrix sensing despite perturbations.

problem Understanding robustness of gradient descent in matrix sensing.
method Developed perturbed gradient flow to capture noise and improve robustness.
result Gradient descent is robust to perturbations in matrix sensing.

We prove the convergence of Kähler-Ricci flow with some small initial curvature conditions. As applications, we discuss the convergence of Kähler-Ricci flow when the complex structure varies on a Kähler-Einstein manifold.

2008-01-20abs ↗pdf ↗

Numerical simulations show stability of Type-II singularities in noncompact hypersurfaces.

problem Stability of Type-II singularities in noncompact hypersurfaces with rotationally-symmetric perturbations.
method Adaptation of the overlap method to include angular dependence.
result MCF of noncompact hypersurfaces with angular dependence behaves similarly to rotationally-symmetric perturbations, developing Type-II or Type-I singularities.

In this paper we consider the polyharmonic heat flow of a closed curve in the plane. Our main result is that closed initial data with initially small normalised oscillation of curvature and isoperimetric defect flows exponentially fast in the C^infty-topology to a simple circle. Our results yield a characterisation of …

2015-05-12abs ↗pdf ↗

Large learning rates lead to optimal generalization if chosen carefully.

problem Understanding the optimal range of large learning rates for neural network training.
method Empirical study focusing on two questions: optimal initial LR range and differences between models trained with different LRs.
result Optimal initial learning rates slightly above the convergence threshold lead to optimal results after fine-tuning with a small LR or weight averaging.

We consider closed immersed hypersurfaces evolving by surface diffusion flow, and perform an analysis based on local and global integral estimates. First we show that a properly immersed stationary (ΔH \equiv 0) hypersurface in \R^3 or \R^4 with restricted growth of the curvature at infinity and small total tracefree c…

2012-05-26abs ↗pdf ↗

In this paper, we study the torsion flow which is served as the CR analogue of the Ricci flow in a closed pseudohermitian manifold. We show that there exists a unique smooth solution to the CR torsion flow in a small time interval with the CR pluriharmonic function as an initial data. In spirit, it is the CR analogue o…

2018-04-18abs ↗pdf ↗

We prove the existence of the flow by curvature of regular planar networks starting from an initial network which is non-regular. The proof relies on a monotonicity formula for expanding solutions and a local regularity result for the network flow in the spirit of B. White's local regularity theorem for mean curvature …

2014-07-17abs ↗pdf ↗

In this article we provide a formulation of empirical bayes described by Atchade (2011) to tune the hyperparameters of priors used in bayesian set up of collaborative filter. We implement the same in MovieLens small dataset. We see that it can be used to get a good initial choice for the parameters. It can also be used…

2017-07-07abs ↗pdf ↗