Proves critical points of ADM mass correspond to specific initial data sets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The principle of convergence stability for geometric flows is the combination of the continuous dependence of the flow on initial conditions, with the stability of fixed points. It implies that if the flow from an initial state exists for all time and converges to a stable fixed point, then the flows of solutions…
We study exponential Levy models with change-point which is a random variable, independent from initial Levy processes. On canonical space with initially enlarged filtration we describe all equivalent martingale measures for change-point model and we give the conditions for the existence of f-divergence minimal equival…
Paper proves new method for constructing initial data in general relativity.
The paper studies neural networks' convergence near origin and saddle points.
We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that under the ZAS initialization, for an arbitrary target matrix, gradient descent con…
The paper analyzes how good initial guesses affect the amount of data needed for low-rank matrix recovery.
We develop a gluing construction which adds scaled and truncated asymptotically Euclidean solutions of the Einstein constraint equations to compact solutions with potentially non-trivial cosmological constants. The result is a one-parameter family of initial data which has ordinary and scaled "point-particle" limits an…
Early training of deep neural networks leads to small, directionally converging weights.
We establish that first-order methods avoid saddle points for almost all initializations. Our results apply to a wide variety of first-order methods, including gradient descent, block coordinate descent, mirror descent and variants thereof. The connecting thread is that such algorithms can be studied from a dynamical s…
New algorithms use outsourced data to improve model training efficiency.
Many modern learning tasks involve fitting nonlinear models to data which are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset. Due to this overparameterization, the training loss may have infinitely many global minima and it is critical to understand the …
Gradient descent with small initialization solves matrix completion without regularization.
In this article, we extend Huisken's theorem that convex surfaces flow to round points by mean curvature flow. We construct certain classes of mean convex and non-mean convex hypersurfaces that shrink to round points and use these constructions to create pathological examples of flows. We find a sequence of flows that …
SGD transitions between maxima and minima with varying time scales.
Study focal surfaces of wave fronts with unbounded curvatures.
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
We present a local gluing construction for general relativistic initial data sets. The method applies to generic initial data, in a sense which is made precise. In particular the trace of the extrinsic curvature is not assumed to be constant near the gluing points, which was the case for previous such constructions. No…
Cosine similarity can force points to grow in magnitude, causing convergence issues.
We develop a general duality between neural networks and compositional kernels, striving towards a better understanding of deep learning. We show that initial representations generated by common random initializations are sufficiently rich to express all functions in the dual kernel space. Hence, though the training ob…
We consider compact convex hypersurfaces contracting by functions of their curvature. Under the mean curvature flow, uniformly convex smooth initial hypersurfaces evolve to remain smooth and uniformly convex, and contract to points after finite time. The same holds if the initial data is only weakly convex or non-smoot…
Develops path integral for spiked tensor model dynamics.
Nonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoi…
We establish an optimal gluing construction for general relativistic initial data sets. The construction is optimal in two distinct ways. First, it applies to generic initial data sets and the required (generically satisfied) hypotheses are geometrically and physically natural. Secondly, the construction is completely …
Let be a compact -dimensional Riemannian manifold with a finite number of singular points, where the metric is asymptotic to a non-negatively curved cone over . We show that there exists a smooth Ricci flow starting from such a metric with curvature decaying like C/t. The initial metr…
Paper refutes EM convergence theory and introduces a new EM algorithm.
The study examines singularities and geometric properties of surfaces derived from frontals with specific singular points.
This article briefly introduced Arthur and Vassilvitshii's work on \textbf{k-means++} algorithm and further generalized the center initialization process. It is found that choosing the most distant sample point from the nearest center as new center can mostly have the same effect as the center initialization process in…
We provide larger step-size restrictions for which gradient descent based algorithms (almost surely) avoid strict saddle points. In particular, consider a twice differentiable (non-convex) objective function whose gradient has Lipschitz constant L and whose Hessian is well-behaved. We prove that the probability of init…
MIK improves t-SNE's local structure preservation in biological sequence data.
Over a compact oriented manifold, the space of Riemannian metrics and normalised positive volume forms admits a natural pseudo-Riemannian metric , which is useful for the study of Perelman's functional. We show that if the initial speed of a -geodesic is -orthogonal to the tangent space to the or…
Many optimization methods for generating black-box adversarial examples have been proposed, but the aspect of initializing said optimizers has not been considered in much detail. We show that the choice of starting points is indeed crucial, and that the performance of state-of-the-art attacks depends on it. First, we d…
Barren plateaus are not an average-case phenomenon, but a highly non-unique problem.
Weak base-point freeness leads to Kähler-Ricci flow diameter bounds.
New method improves convergence of spatial filters in neural networks.
Let be the scattering relation on a compact Riemannian manifold with non-necessarily convex boundary, that maps initial points of geodesic rays on the boundary and initial directions to the outgoing point on the boundary and the outgoing direction. Let be the length of that geodesic ray. We study the que…
In label-noise learning, \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \textit{anchor points} (i.e…
New method for robust fixed-point smoothing without state augmentation.
Unified framework detects changes in complex system models.
We construct large families of initial data sets for the vacuum Einstein equations with positive cosmological constant which contain exactly Delaunay ends; these are non-trivial initial data sets which coincide with those for the Kottler-Schwarzschild-de Sitter metrics in regions of infinite extent. From the purely Rie…
DEQs converge to optimal solutions with mild over-parameterization.
New method avoids spurious critical points for low-rank matrix recovery.
The paper studies how curves evolve under area constraints and converges to a critical point.
We prove local existence for the second order Renormalization Group flow initial value problem on closed Riemannian manifolds in general dimensions, for initial metrics whose sectional curvatures satisfy the condition , at all points and planes . This extends results…
In this paper, we introduce a new parabolic equation on Kähler manifolds. The static point of this flow is related to the existence of a lower bound of the Mabuchi energy. In this paper, we prove the flow always exists for all times for any initial smooth data. Further more, if the initial metric has non-negative bisec…
Parametric models, and particularly neural networks, require weight initialization as a starting point for gradient-based optimization. Recent work shows that a specific initial parameter set can be learned from a population of supervised learning tasks. Using this initial parameter set enables a fast convergence for u…
Poor (even random) starting points for learning/training/optimization are common in machine learning. In many settings, the method of Robbins and Monro (online stochastic gradient descent) is known to be optimal for good starting points, but may not be optimal for poor starting points -- indeed, for poor starting point…
On the space of positive 3-forms on a seven-manifold, we study a natural functional whose critical points induce metrics with holonomy contained in . We prove short-time existence and uniqueness for its negative gradient flow. Furthermore, we show that the flow exists for all times and converges modulo diffeomorph…