Gradient descent converges globally in deep linear residual networks with ZAS initialization.
problem Optimizing deep linear residual networks for convergence.
method Zero-asymmetric (ZAS) initialization for gradient descent.
result Gradient descent converges to an ε-optimal point in O(L^3 log(1/ε)) iterations.
Unique global solutions found for specific initial data.
problem Einstein-scalar-field equations with specific initial conditions.
method Spherically symmetric analysis of small, slowly decaying data.
result Unique global solutions exist for the equations.
Global existence of Yamabe flow on non-compact manifolds with unbounded initial curvature.
problem Global existence of Yamabe flow on non-compact manifolds with unbounded initial curvature.
method Assumption of conformally equivalent initial metric to a complete background metric with bounded scalar curvature and positive Yamabe invariant.
result Global existence of Yamabe flow without requiring initial curvature bounds.
Proves stability of Minkowski space for specific initial data.
problem Stability of Minkowski space under spacelike-characteristic initial data.
method Vectorfield method and bootstrapping argument, with new geometric constructions.
result Global nonlinear stability of Minkowski space proved for the spacelike-characteristic Cauchy problem.
Gradient descent solves non-convex neural networks with random initialization.
problem Gradient descent can solve non-convex neural networks with random initialization.
method Gradient descent, over-parameterized neural networks, random initialization, strong convexity-like property.
result Gradient descent converges to a globally optimal solution at a linear rate.
Gradient descent proves global convergence for 4-layer matrix factorization.
problem Global convergence of gradient descent on four-layer matrix factorization under random initialization.
method New techniques to show saddle-avoidance properties and extend eigenvalue theories.
result Polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization.
We show the existence of a global unique and analytic solution for the mean curvature flow, the surface diffusion flow and the Willmore flow of entire graphs for Lipschitz initial data with small Lipschitz norm. We also show the existence of a global unique and analytic solution to the Ricci-DeTurck flow on euclidean s…
Global implicit function theorem for Fréchet spaces, solving derivative loss problems.
problem Solving initial value problems with derivative loss in Fréchet spaces.
method Global implicit function theorems for Keller's Cc1-mappings in Fréchet spaces, applied through submersions and transversality. result Global existence and uniqueness of solutions to initial value problems with derivative loss.
Global existence of Yamabe flows on hyperbolic space proved without curvature bounds.
problem Global existence of Yamabe flows on hyperbolic space without completeness or curvature bounds.
method Instantaneously complete initial metrics, no curvature bounds required.
result Global existence of Yamabe flows on hyperbolic space of arbitrary dimension m≥3. We present a local gluing construction for general relativistic initial data sets. The method applies to generic initial data, in a sense which is made precise. In particular the trace of the extrinsic curvature is not assumed to be constant near the gluing points, which was the case for previous such constructions. No…
The paper proves rigidity results for compact initial data sets.
problem Understanding the structure of compact initial data sets.
method Proving rigidity results under specific conditions.
result Global version of the main result in [15] is obtained.
Neural networks with Xavier initialization converge to global minimum in the scaling limit.
problem Optimizing neural networks with Xavier initialization in the large network limit.
method Stochastic analysis and convergence to a random ODE with a Gaussian distribution.
result The neural network converges to a global minimum in the limit, with zero loss.
New method creates vacuum data at minimal and borderline decay thresholds.
problem Creating vacuum initial data at specific decay thresholds.
method Conical solution-operator method applied to vacuum asymptotically flat initial data.
result Demonstrates global and exterior stability of Minkowski spacetime.
Global convergence of multilayer neural networks proven for any depth.
problem Global convergence of multilayer neural networks in the mean field regime.
method Mean field limit framework, neuronal embedding, bidirectional diversity condition.
result Global convergence for multilayer networks of any depths, including correlated initializations.
Proves global existence and uniqueness of solutions for Einstein-scalar-field equations.
problem Global existence and uniqueness of solutions for specific Einstein-scalar-field equations.
method Proves global existence and uniqueness of classical solutions with small initial data and wake-like decaying null infinity.
result Global existence and uniqueness of solutions for the equations with wake-like decaying null infinity.
Paper refutes EM convergence theory and introduces a new EM algorithm.
problem The convergence theory of the EM algorithm is incorrect and affects its performance.
method Proposes a new EM algorithm called the Channel Matching (CM) EM algorithm and provides an initialization map.
result The locally maximal Q can affect the convergent speed but not the global convergence.
In this note a proof is given for global existence and uniqueness of minimal surfaces of Lorentzian type from a cylinder into globally hyperbolic Lorentzian manifolds for given initial values up to the first derivatives.
Proposes a new method to initialize neural networks by estimating global curvature of weights.
problem Improving the initialization of neural networks for better training and convergence.
method Estimates the global curvature of weights across layers using the Hessian matrix norm.
result The proposed method helps in more rigorously initializing weights, leading to better performance.
Gradient descent finds the shortest path to global optima in overparameterized models.
problem Training overparameterized nonlinear models with gradient descent.
method Gradient descent, martingale techniques.
result Gradient descent converges geometrically to a global optimum near the initial point.
Solves initial boundary value problem for vacuum Einstein equations and proves geometric uniqueness.
problem Initial boundary value problem for vacuum Einstein equations.
method Formulated IBVP, solved simultaneously in local harmonic coordinates, constructed unique maximal globally hyperbolic solution.
result Vacuum spacetimes satisfying fixed initial-boundary conditions and corner conditions are geometrically unique near the initial surface.
Unique solutions found for wave-like decaying null infinity equations.
problem Wave-like decaying null infinity equations with spherically symmetric Einstein-scalar-field.
method Local and global unique solutions for small initial data.
result Sharp decaying condition for unique solutions.
XGL uses global explanations to guide human supervision in machine learning.
problem Improving model quality through human-machine interaction.
method XGL employs global explanations to guide human selection of informative examples.
result XGL avoids overselling the model's quality and performs comparably to other strategies.
Solves the Cauchy problem for linearised Einstein equation on globally hyperbolic spacetimes.
problem Initial value problem for gravitational waves on globally hyperbolic vacuum spacetimes.
method Proves the solution map is an isomorphism of locally convex topological vector spaces and solves linearised constraint equations on closed manifolds.
result Well-posedness of the Cauchy problem for gravitational waves on globally hyperbolic spacetimes.
Study shows global oscillatory solutions for Yang-Mills heat flow in 4D space.
problem Investigating long-time dynamics of Yang-Mills heat flow with specific initial data.
method Analysis of SO(4)-equivariant Yang-Mills heat flow with SU(2) group in 4D space. result Global solutions can exhibit oscillatory behavior at time infinity.
Under weak regularity assumptions, only, we develop a fully geometric theory of vacuum Einstein spacetimes with T2 symmetry, establish the global well-posedness of the initial value problem for Einstein's field equations, and investigate the global causal structure of the constructed spacetimes. Our weak regularity ass…
We investigate the initial value problem for the Einstein-Euler equations of general relativity under the assumption of Gowdy symmetry on T3, and we construct matter spacetimes with low regularity. These spacetimes admit, both, impulsive gravitational waves in the metric (for instance, Dirac mass curvature singularitie…
Proves global existence of maps with curvature term on expanding spacetimes.
problem Global existence of Dirac-wave maps with curvature term on expanding spacetimes.
method Proves global existence with small initial data on globally hyperbolic manifolds with growth condition.
result Global existence of maps proved for small initial data.
Global properties of maximal future Cauchy developments of stationary, m-dimensional asymptotically flat initial data with an outer trapped boundary are analyzed. We prove that, whenever the matter model is well posed and satisfies the null energy condition, the future Cauchy development of the data is a black hole spa…
New method avoids spurious critical points for low-rank matrix recovery.
problem Low-rank matrix recovery problems on Riemannian manifold.
method Riemannian gradient descent with random initialization.
result Riemannian gradient descent avoids spurious critical points and converges nearly linearly.
AMP method reconstructs rank-one matrices from noisy data efficiently.
problem Reconstructing rank-one matrices with prior structural information from noisy observations.
method Approximate Message Passing (AMP) with random initialization.
result AMP from random initialization converges rapidly and globally.
Gradient descent optimizes deep ReLU networks with proper initialization.
problem Training deep neural networks with ReLU activation.
method Gradient descent and stochastic gradient descent with proper random weight initialization.
result Gradient descent finds global minima for over-parameterized deep ReLU networks.
Gradient descent with random initialization solves phase retrieval problems efficiently.
problem Solving systems of quadratic equations for phase retrieval.
method Gradient descent with random initialization for nonconvex least squares problem.
result Gradient descent achieves near-optimal computational and sample complexities for phase retrieval.
Study well-posedness of SPDE on Riemannian manifolds with rough initial conditions.
problem Well-posedness of parabolic Anderson model on Riemannian manifolds with rough initial conditions.
method Construct intrinsic Gaussian noises, explore global geometry, use Feynman-Kac formula.
result Show well-posedness with non-positive curvature and conditions on α. We establish both local and global well-posedness for the heat flow of polyharmonic maps from Rn to a compact Riemannian manifold without boundary for initial data with small BMO norms.
New method for symmetric matrix completion using ReLU sampling.
problem Symmetric positive semi-definite low-rank matrix completion with deterministic entry-dependent sampling.
method ReLU sampling, gradient descent with tailored initialization.
result Gradient descent with tailored initialization achieves global minima.
Global existence and convergence of heat flow for p-harmonic maps.
problem Global existence and convergence of heat flow for p-harmonic maps between manifolds.
method Analysis of heat flow equations for p-harmonic maps.
result Global existence and convergence of heat flow for p-harmonic maps under certain conditions.
Study curve flows with global forcing terms using a distance comparison principle.
problem Analyse the behavior of curves under curve flows with global forcing terms.
method Prove a distance comparison principle for curve shortening flow with arbitrary global forcing terms.
result Established a distance comparison principle for curve flows with global forcing terms.
We give a global description of envelopes of geodesic tangents of regular curves in (not necessarily convex) Riemannian surfaces. We prove that such an envelope is the union of the curve itself, its inflectional geodesics and its tangential caustics (formed by the conjugate points to those of the initial curve along th…
DLNs dynamics change with variance, leading to saddle-to-saddle training phases.
problem Understanding the dynamics of DLNs with varying initialization variance.
method Analyzing the phase transition of DLNs' dynamics as variance changes.
result Gradient descent visits a sequence of saddles, reaching a sparse global minimum.
EM algorithm converges globally for two-component mixed linear regression.
problem Global convergence of EM algorithm for mixed linear regression.
method Developed new theoretical analysis for EM algorithm convergence in mixed linear regression.
result EM algorithm converges globally for two-component mixed linear regression.
Global convergence for robust regression problems via IRLS with enhancements.
problem Global convergence for robust regression problems.
method Augmentations to IRLS to ensure global recovery and improved robustness.
result Global recovery guarantees for robust regression problems, outperforming state-of-the-art algorithms.
We show that there are no spurious local minima in the non-convex factorized parametrization of low-rank matrix recovery from incoherent linear measurements. With noisy measurements we show all local minima are very close to a global optimum. Together with a curvature bound at saddle points, this yields a polynomial ti…
SGD batch size affects autoencoder global minima sparsity and sharpness.
problem Investigating how batch size impacts autoencoder learning.
method Non-convex autoencoder training with SGD, varying batch sizes.
result SGD batch size influences global minimum sparsity and sharpness.
We establish the global well-posedness of the initial value problem for the Schrodinger map flow for maps from the real line into Kahler manifolds and for maps from the circle into Riemann surfaces. This partially resolves a conjecture of W.-Y. Ding.
This paper shows how deep neural networks can learn rich, independent features that significantly deviate from initialization.
problem Understanding how deep neural networks achieve meaningful feature learning and global convergence.
method Investigation of infinitely wide, L-layer neural networks using the tensor program framework under Maximal Update parametrization. result SGD enables these networks to learn linearly independent features that substantially deviate from their initial values, capturing relevant data information.
Willmore flow converges globally for surfaces with rotational symmetry below a specific energy threshold.
problem Global existence and convergence of Willmore flow with Dirichlet boundary conditions.
method Considered surfaces with rotational symmetry, proved global existence and convergence for initial data below a sharp energy threshold.
result Sharp threshold for global existence and convergence of Willmore flow depends on boundary conditions.
We study the global theory of linear wave equations for sections of vector bundles over globally hyperbolic Lorentz manifolds. We introduce spaces of finite energy sections and show well-posedness of the Cauchy problem in those spaces. These spaces depend in general on the choice of a time function but it turns out tha…
This paper studies transformer learning dynamics and initialization.
problem Understanding how transformers learn Markov chains and the role of initialization.
method First-order Markov chains and single-layer transformers, proving learning dynamics and conditions for convergence.
result Transformer parameters can converge to global or local minima based on initialization and Markovian data properties.