Study shows momentum-based optimizers like Muon and MomentumGD bias towards KKT points in smooth homogeneous models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work analyzes the maximum-margin bias in quasi-homogeneous neural networks.
In this paper we investigate the strict convexity and the differentiability properties of the stable norm, which corresponds to the homogenized surface tension for a periodic perimeter homogenization problem (in a regular and uniformly elliptic case). We prove that it is always differentiable in totally irrational dire…
Researchers classify special curved spheres in a complex space.
There is considered the problem of describing up to linear conformal equivalence those harmonic cubic homogeneous polynomials for which the squared-norm of the Hessian is a nonzero multiple of the quadratic form defining the Euclidean metric. Solutions are constructed in all dimensions and solutions are classified in d…
We study the deformation of the three-dimensional conformal structures by the Ricci flow. We drive the evolution equation of Cotton-York tensor and the L1-norm of it under the Ricci flow. In particular, we investigate the behavior of the L1-norm of the Cotton-York tensor under the Ricci flow on three-dimensional simply…
Carnot groups can be polarized if they have specific coordinate systems.
A unique Kähler potential on the unit ball is identified with constant differential norm.
Sharp stability of isometries on Heisenberg group proven.
With an eye toward understanding complexity control in deep learning, we study how infinitesimal regularization or gradient descent optimization lead to margin maximizing solutions in both homogeneous and non-homogeneous models, extending previous work that focused on infinitesimal regularization only in homogeneous mo…
The real homology of a compact, n-dimensional Riemannian manifold M is naturally endowed with the stable norm. The stable norm of a homology class is the minimal Riemannian volume of its representatives. If M is orientable the stable norm on H_{n-1}(M,R) is a homogenized version of the Riemannian (n-1)-volume. We study…
We review and extend here some recent results on the existence of minimal surfaces and isoperimetric sets in non homogeneous and anisotropic periodic media. We also describe the qualitative properties of the homogenized surface tension, also known as stable norm (or minimal action) in Weak KAM theory. In particular we …
New capacity measure for deep ReLU networks derived from weight norms.
New regularizer improves neural network robustness and generalization.
GD iterates for non-homogeneous deep nets increase margin and converge in direction.
We prove uniform curvature estimates for homogeneous Ricci flows: For a solution defined on the norm of the curvature tensor at time is bounded by the maximum of and . This is used to show that solutions with finite extinction time are Type I, immortal solutions ar…
Unified theory of deep neural networks with diverse activations.
The real homology of a compact Riemannian manifold is naturally endowed with the stable norm. The stable norm on arises from the Riemannian length functional by homogenization. It is difficult and interesting to decide which norms on the finite-dimensional vector space are st…
Introduces a natural parallel translation for navigation data.
In this paper, we give a new generalization error bound of Multiple Kernel Learning (MKL) for a general class of regularizations, and discuss what kind of regularization gives a favorable predictive accuracy. Our main target in this paper is dense type regularizations including \ellp-MKL. According to the recent numeri…
In this dissertation we propose alternative analysis of distributed stochastic gradient descent (SGD) algorithms that rely on spectral properties of the data covariance. As a consequence we can relate questions pertaining to speedups and convergence rates for distributed SGD to the data distribution instead of the regu…
We study biinvariant word metrics on groups. We provide an efficient algorithm for computing the biinvariant word norm on a finitely generated free group and we construct an isometric embedding of a locally compact tree into the biinvariant Cayley graph of a nonabelian free group. We investigate the geometry of cyclic …
Let and be two Kähler manifolds. One says that and are "relatives" if they share a non-trivial Kähler submanifold , namely, if there exist two holomorphic and isometric immersions (Kähler immersions) and . In this paper we show that a bounded homogeneous domain …
Early training of deep neural networks leads to small, directionally converging weights.
We use the bracket flow/algebraic soliton approach to study the Laplacian flow of -structures and its solitons in the homogeneous case. We prove that any homogeneous Laplacian soliton is equivalent to a semi-algebraic soliton (i.e.\ a -invariant -structure on a homogeneous space that flows by pull-ba…
Infinite diameter proved for contractible loops space.
The study introduces new tensors for almost Finsler manifolds and analyzes their properties.
The study shows how to regularize weakly harmonic maps using Sobolev norms and Coulomb frames.
In recent years, there has been a growing interest in geometric evolution in heterogeneous media. Here we consider curvature driven fows of planar curves, with an additional space-dependent forcing term. Motivated by a homogenization problem, we look for estimates which depend only on the uniform norm of the forcing te…
Unified framework for estimating high-dimensional conditional factor models.
Let F be the fundamental group of S, where S is a compact, connected, oriented surface with negative Euler characteristic and nonempty boundary. (1) The projective class of the chain \partial S in B_1(F) intersects the interior of a codimension one face of the unit ball in the stable commutator length pseudo-norm. (2) …
Study on geodesic distances on SE(3)/SO(2) in machine learning.
Study of unitary and groupoid orbits of normal operators, focusing on manifold structures and spectral conditions.
A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effect of finding the minimum RKHS norm solution. This stands in contrast to other studies which demonstr…
Bounded symmetric domains are biholomorphic to tube domains over Finsler symmetric cones.
We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural networks with linear, ReLU or Leaky ReLU activation. We rigorously prove that gradient flow (i.e. gradient descent with infinitesimal step …
Gradient descent is a simple and widely used optimization method for machine learning. For homogeneous linear classifiers applied to separable data, gradient descent has been shown to converge to the maximal margin (or equivalently, the minimal norm) solution for various smooth loss functions. The previous theory does …
The notion of nonpositive curvature in Alexandrov's sense is extended to include p-uniformly convex Banach spaces. Infinite dimensional manifolds of semi-negative curvature with a p-uniformly convex tangent norm fall in this class on nonpositively curved spaces, and several well-known results, such as existence and uni…
The paper studies the asymptotic behavior of HCMA equations on ALE Kahler manifolds.
The paper proves lifting theorems for complex representations of finite groups.
Catapult phase in neural nets shows exponential loss growth before quick decrease.
A helical CR structure is a decomposition of a real Euclidean space into an even-dimensional horizontal subspace and its orthogonal vertical complement, together with an almost complex structure on the horizontal space and a marked vector in the vertical space. We prove an equivalence between such structures and step t…
A new algorithm trains deep neural networks by adding neurons greedily.
We prove an estimate for solutions to the linearized Ricci flow system on closed 3-manifolds. This estimate is a generalization of Hamilton's pinching is preserved estimate for the Ricci curvatures of solutions to the Ricci flow on 3-manifolds with positive Ricci curvature. In our estimate we make no assumption on the …
Study on how initialization scale affects neural network training regimes.
The success of deep convolutional architectures is often attributed in part to their ability to learn multiscale and invariant representations of natural signals. However, a precise study of these properties and how they affect learning guarantees is still missing. In this paper, we consider deep convolutional represen…
Introduces new types of homogeneous spaces and their properties.
Study improves estimation of functions from noisy data using convex penalties.