Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
problem Conditions for Willmore surfaces to have finite ends or finite total curvature.
method Analyzes scale-invariant second fundamental form near infinity.
result Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …
Using the monotonicity formulas of Colding and Minicozzi, we prove that on any complete, non-parabolic Riemannian manifold (M3,g) with non-negative Ricci curvature, the asymptotic weighted scaling invariant integral of scalar curvature has an explicit bound in form of asymptotic volume ratio.
Scales attention for long contexts in LLMs.
problem Development of attention mechanisms for long context inference.
method Scale-invariant total attention and sparsity conditions, with a position-dependent transformation of logits.
result Scale-invariant attention scheme improves validation loss and long-context retrieval.
Three training regimes found for scale-invariant neural networks on the sphere.
problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
Proves uniqueness of Ricci flow with scaling invariant estimates.
problem Proving uniqueness of Ricci flow with scaling invariant curvature bound.
method Solving Ricci-harmonic map heat flow in unbounded curvature background.
result Complete Ricci flow starting from uniformly non-collapsed, non-negatively curved manifold is unique in dimension three.
Critical points of scale-invariant curvature energies in 4D are analytic.
problem Analyzing critical points of curvature energies in 4D manifolds.
method Applying Noether's theorem to identify conservation laws and lower order elliptic system of PDEs, then using integrability by compensation and interpolation theory.
result Critical points of scale-invariant curvature energies in 4D are analytic.
It is well known that neural networks with rectified linear units (ReLU) activation functions are positively scale-invariant. Conventional algorithms like stochastic gradient descent optimize the neural networks in the vector space of weights, which is, however, not positively scale-invariant. This mismatch may lead to…
New bounds on self-normalized martingales improve online linear regression performance.
problem Improving regret bounds in online linear regression.
method Characterizing scale-invariant bounds on self-normalized martingales.
result For d=1, O(logT) doubly-uniform regret is possible; for d>1, sublinear doubly-uniform regret is impossible. In order to study the application of artificial intelligence (AI) to dental imaging, we applied AI technology to classify a set of panoramic radiographs using (a) a convolutional neural network (CNN) which is a form of an artificial neural network (ANN), (b) representative image cognition algorithms that implement scal…
We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
problem Premature decay of effective step sizes in momentum-based optimizers for scale-invariant weights.
method Proposes SGDP and AdamP to eliminate the radial component at each optimizer step, preserving convergence properties.
result Uniform gains across multiple benchmarks, improving model performance.
Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.
problem Finding and characterizing minimizers and critical points of scale-invariant tangent-point energies for closed curves.
method Develops convergence and regularity theories based on fractional Sobolev spaces and new energy functionals.
result Minimizing sequences converge to locally critical embeddings in all but finitely many points, and locally critical embeddings are regular.
We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…
Study shows how near crushing singularities, Kasner-like regions can exist.
problem Understanding spatial volume densities near crushing singularities.
method Relates existence of Kasner-like regions to asymptotics of spatial volume densities under scale-invariant curvature bounds.
result Kasner-like regions can exist near crushing singularities under certain curvature conditions.
New learning dynamics achieve fast convergence in games without needing to know utility scales.
problem Fast convergence guarantees in learning games require prior knowledge of utility scales.
method Developed scale-free and scale-invariant learning dynamics using optimistic follow-the-regularized-leader with adaptive learning rates and clipping techniques.
result Achieved fast convergence rates to Nash and correlated equilibria without prior utility scale knowledge.
Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.
problem Why Adam performs better with β1=β2. method Formalized gradient scale invariance and proved it for Adam with equal β1 and β2. result Adam becomes gradient scale invariant of first order if and only if β1=β2. L-CNNs approximate gauge actions, revealing fixed points with no lattice artifacts.
problem Approximating gauge actions with lattice artifacts.
method Lattice gauge-equivariant convolutional neural networks (L-CNNs).
result L-CNNs provide fixed point actions with no lattice artifacts.
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.
This paper extends compositional data analysis using graph signal processing.
problem Traditional log-ratios between all variables are not suitable for specific variable relationships.
method Linking compositional data analysis with graph signal processing, it considers only selected log-ratios.
result The approach retains desirable properties of scale invariance and compositional coherence.
We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…
ASAM improves deep neural network generalization by adapting sharpness to scale.
problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.
Unified framework for scale-invariant representation learning using MAPCA.
problem Learning invariant representations in data.
method Metric-Aware Principal Component Analysis (MAPCA) based on generalized eigenproblem.
result MAPCA provides a unified geometric language for various self-supervised learning objectives.
The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.
problem Nonlinear dynamics of minimal resistance in radial fields.
method Analysis of two non-equilibrium scenarios: scale-invariant free expansion and incompressible source flow.
result Incompressible flow acts as a structural regularizer, admitting unique, smooth, and strictly concave solutions.
New method trains deep networks robustly without adaptive methods.
problem Training deep networks with robustness and efficiency.
method Scale invariant architecture + SGD + weight decay + gradient clipping.
result SGD can achieve similar performance to adaptive methods like Adam.
Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.
problem Proving long-time existence and topological rigidity for manifolds with pinched scale-invariant integral curvature.
method Proves long-time existence of Ricci flow for manifolds with bounded curvature and pinched scale-invariant integral curvature, converging to a flat metric.
result Flow converges to a flat metric, implying topological rigidity of the manifold.
New approach improves classification guarantees by focusing on direction rather than regression risk.
problem Improving classification guarantees in binary classification problems.
method Establishing a geometric distinction between classification and regression, leveraging scale invariance.
result Improved guarantees for classification risk compared to regression risk.
We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…
In seeking for sparse and efficient neural network models, many previous works investigated on enforcing L1 or L0 regularizers to encourage weight sparsity during training. The L0 regularizer measures the parameter sparsity directly and is invariant to the scaling of parameter values, but it cannot provide useful gradi…
Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
problem Modeling non-Gaussian time-series with stationary increments.
method Complex wavelet transform for scale variations, joint correlation matrix for scale dependencies, second wavelet transform for diagonalization, maximum entropy models conditioned by scattering spectra coefficients.
result Scattering spectra of self-similar processes are scale invariant, allowing statistical testing and generation of new time-series.
Proves planarity and convexity for ancient solutions of mean curvature flow.
problem Ancient solutions of mean curvature flow in higher codimension.
method Parabolically scale-invariant variation of planarity estimate, convexity proof for pinched solutions.
result Characterizes certain pinched complete ancient solutions and shrinkers in higher codimension.
NOTEARS fails to identify true causal relationships from data.
problem Identifying true causal relationships from observational data.
method NOTEARS aims to identify a parsimonious DAG from data explaining residual variance.
result NOTEARS is not suitable for identifying true causal relationships.
The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a R2 term) are applied to explain the personal income distribution. A scale invariance exists if there is not the R2 term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…
Enhances power of covariance matrix tests for high-dimensional data.
problem Testing large covariance matrices in high-dimensional data.
method Proposes a new Fisher's combined probability test for quadratic form and maximum form statistics.
result Boosts power against more general alternatives.
We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after unspecified marginally monotone transformations, the distributions are multivariate Gaussian. COCA improves upon PCA and sparse PCA in three as…
Alternative proof and extension of curvature estimates for minimal immersions.
problem Curvature estimates and Bernstein-type theorems for minimal immersions.
method Iteration method à la De Giorgi, ε-regularity theorem, Caccioppoli inequalities.
result Extension of Schoen--Simon--Yau and Schoen--Simon theorems to 6-dimensional stable minimal immersions.
The statistical properties of the multipliers of the absolute returns are investigated using one-minute high-frequency data of financial time series. The multiplier distribution is found to be independent of the box size s when s is larger than some crossover scale, providing direct evidence of the existence of sca…
We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…
Consider a family of smooth immersions F(⋅,t):Mn→Rn+1 of closed hypersurfaces in Rn+1 moving by the mean curvature flow ∂t∂F(p,t)=−H(p,t)⋅ν(p,t), for t∈[0,T). In \cite{Cooper} Cooper has recently proved that the mean curvature blows up at the s…
This article investigates stationary surfaces with boundaries, which arise as the critical points of functionals dependent on curvature. Precisely, a generalized "bending energy" functional W is considered which involves a Lagrangian that is symmetric in the principal curvatures. The first variation of $\ma…
Reduces symplectic Hamiltonian systems to contact systems, realizing Poincaré's dream.
problem Scaling symmetries in Hamiltonian systems.
method Contact reduction of symplectic Hamiltonian systems.
result Generically possible reduction to contact Hamiltonian systems, reducing inputs needed.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
The paper extends a Harnack inequality to noncompact evolving hypersurfaces.
problem Proving a Harnack inequality for noncompact evolving hypersurfaces.
method Using a differential Harnack inequality for noncompact convex hypersurfaces flowing with normal speed based on their principal curvatures.
result The extension of Andrews' result to noncompact hypersurfaces.
Deep learning models are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on benign inputs. However, under the black-box setting, most existing adversaries often have a poor transferability to attack other defense models. In this work, from the perspective of regarding the advers…
Consider a family of smooth immersions F(⋅,t):Mn→Rn+1 of closed hypersurfaces in Rn+1 moving by the mean curvature flow ∂t∂F(p,t)=−H(p,t)⋅ν(p,t), for t∈[0,T). We show that at the first singular time of the mean curvature flow, certain subcritic…
We consider inverse curvature flows in hyperbolic space with starshaped initial hypersurface, driven by positive powers of a homogeneous curvature function. The solutions exist for all time and, after rescaling, converge to a sphere.
Generalized algorithm for translation and scale-invariant prediction.
problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.