Proposes a method to compare noisy high-dimensional datasets with low-dimensional manifolds.
problem Comparing distributions on manifolds in noisy high-dimensional datasets.
method Linking low-rank structure to manifold geometry, developing a scale-invariant distance measure.
result Superior robustness and statistical power compared to existing methods.
Proposes GFMMD for comparing signals on graphs.
problem Computing distances between distributions on graphs.
method Graph Fourier MMD (GFMMD) using optimal witness functions.
result Analytical solution and embedding of distributions.
A new method detects small holes in noisy data.
problem Detecting small holes in high-density regions from noise.
method Robust Density-Aware Distance (RDAD) filtration, incorporating distance-to-measure concept.
result The RDAD filtration prolongs the persistences of small holes, making them distinguishable from noise.
Scales attention for long contexts in LLMs.
problem Development of attention mechanisms for long context inference.
method Scale-invariant total attention and sparsity conditions, with a position-dependent transformation of logits.
result Scale-invariant attention scheme improves validation loss and long-context retrieval.
Three training regimes found for scale-invariant neural networks on the sphere.
problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
Proves uniqueness of Ricci flow with scaling invariant estimates.
problem Proving uniqueness of Ricci flow with scaling invariant curvature bound.
method Solving Ricci-harmonic map heat flow in unbounded curvature background.
result Complete Ricci flow starting from uniformly non-collapsed, non-negatively curved manifold is unique in dimension three.
Critical points of scale-invariant curvature energies in 4D are analytic.
problem Analyzing critical points of curvature energies in 4D manifolds.
method Applying Noether's theorem to identify conservation laws and lower order elliptic system of PDEs, then using integrability by compensation and interpolation theory.
result Critical points of scale-invariant curvature energies in 4D are analytic.
New bounds on self-normalized martingales improve online linear regression performance.
problem Improving regret bounds in online linear regression.
method Characterizing scale-invariant bounds on self-normalized martingales.
result For d=1, O(logT) doubly-uniform regret is possible; for d>1, sublinear doubly-uniform regret is impossible. We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
problem Premature decay of effective step sizes in momentum-based optimizers for scale-invariant weights.
method Proposes SGDP and AdamP to eliminate the radial component at each optimizer step, preserving convergence properties.
result Uniform gains across multiple benchmarks, improving model performance.
Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
problem Conditions for Willmore surfaces to have finite ends or finite total curvature.
method Analyzes scale-invariant second fundamental form near infinity.
result Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …
Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.
problem Finding and characterizing minimizers and critical points of scale-invariant tangent-point energies for closed curves.
method Develops convergence and regularity theories based on fractional Sobolev spaces and new energy functionals.
result Minimizing sequences converge to locally critical embeddings in all but finitely many points, and locally critical embeddings are regular.
We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…
Study shows how near crushing singularities, Kasner-like regions can exist.
problem Understanding spatial volume densities near crushing singularities.
method Relates existence of Kasner-like regions to asymptotics of spatial volume densities under scale-invariant curvature bounds.
result Kasner-like regions can exist near crushing singularities under certain curvature conditions.
New learning dynamics achieve fast convergence in games without needing to know utility scales.
problem Fast convergence guarantees in learning games require prior knowledge of utility scales.
method Developed scale-free and scale-invariant learning dynamics using optimistic follow-the-regularized-leader with adaptive learning rates and clipping techniques.
result Achieved fast convergence rates to Nash and correlated equilibria without prior utility scale knowledge.
Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.
problem Why Adam performs better with β1=β2. method Formalized gradient scale invariance and proved it for Adam with equal β1 and β2. result Adam becomes gradient scale invariant of first order if and only if β1=β2. A new method treats all variables equally in fitting data.
problem Fitting relationships to data with multiple variables, especially when dependent and independent variables are not clearly defined.
method A general method treating all variables impartially, using geometric mean functional relationships and correlation.
result The method provides coefficients that are easily calculated from covariances or correlations, making it scale-invariant and applicable to various units.
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.
We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…
ASAM improves deep neural network generalization by adapting sharpness to scale.
problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.
This paper analyzes implicit bias in Deep Linear Discriminant Analysis.
problem The implicit bias of Deep Linear Discriminant Analysis.
method Analyzing gradient flow on a L-layer diagonal linear network.
result Under balanced initialization, the network transforms additive updates into multiplicative updates, conserving the (2/L) quasi-norm.
Unified framework for scale-invariant representation learning using MAPCA.
problem Learning invariant representations in data.
method Metric-Aware Principal Component Analysis (MAPCA) based on generalized eigenproblem.
result MAPCA provides a unified geometric language for various self-supervised learning objectives.
The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.
problem Nonlinear dynamics of minimal resistance in radial fields.
method Analysis of two non-equilibrium scenarios: scale-invariant free expansion and incompressible source flow.
result Incompressible flow acts as a structural regularizer, admitting unique, smooth, and strictly concave solutions.
The effectiveness of Convolutional Neural Networks (CNNs) has been substantially attributed to their built-in property of translation equivariance. However, CNNs do not have embedded mechanisms to handle other types of transformations. In this work, we pay attention to scale changes, which regularly appear in various t…
New method trains deep networks robustly without adaptive methods.
problem Training deep networks with robustness and efficiency.
method Scale invariant architecture + SGD + weight decay + gradient clipping.
result SGD can achieve similar performance to adaptive methods like Adam.
Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.
problem Proving long-time existence and topological rigidity for manifolds with pinched scale-invariant integral curvature.
method Proves long-time existence of Ricci flow for manifolds with bounded curvature and pinched scale-invariant integral curvature, converging to a flat metric.
result Flow converges to a flat metric, implying topological rigidity of the manifold.
New approach improves classification guarantees by focusing on direction rather than regression risk.
problem Improving classification guarantees in binary classification problems.
method Establishing a geometric distinction between classification and regression, leveraging scale invariance.
result Improved guarantees for classification risk compared to regression risk.
We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…
New insights show embedding lengths correlate with semantic properties.
problem Contrastive embedding norms ignore embedding magnitudes but correlate with semantic properties.
method Formal theoretical framework and analysis of optimization dynamics.
result Embedding lengths encode semantic information as a byproduct of training.
In seeking for sparse and efficient neural network models, many previous works investigated on enforcing L1 or L0 regularizers to encourage weight sparsity during training. The L0 regularizer measures the parameter sparsity directly and is invariant to the scaling of parameter values, but it cannot provide useful gradi…
Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
problem Modeling non-Gaussian time-series with stationary increments.
method Complex wavelet transform for scale variations, joint correlation matrix for scale dependencies, second wavelet transform for diagonalization, maximum entropy models conditioned by scattering spectra coefficients.
result Scattering spectra of self-similar processes are scale invariant, allowing statistical testing and generation of new time-series.
It is well known that neural networks with rectified linear units (ReLU) activation functions are positively scale-invariant. Conventional algorithms like stochastic gradient descent optimize the neural networks in the vector space of weights, which is, however, not positively scale-invariant. This mismatch may lead to…
Proves planarity and convexity for ancient solutions of mean curvature flow.
problem Ancient solutions of mean curvature flow in higher codimension.
method Parabolically scale-invariant variation of planarity estimate, convexity proof for pinched solutions.
result Characterizes certain pinched complete ancient solutions and shrinkers in higher codimension.
NOTEARS fails to identify true causal relationships from data.
problem Identifying true causal relationships from observational data.
method NOTEARS aims to identify a parsimonious DAG from data explaining residual variance.
result NOTEARS is not suitable for identifying true causal relationships.
The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a R2 term) are applied to explain the personal income distribution. A scale invariance exists if there is not the R2 term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…
We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after unspecified marginally monotone transformations, the distributions are multivariate Gaussian. COCA improves upon PCA and sparse PCA in three as…
The statistical properties of the multipliers of the absolute returns are investigated using one-minute high-frequency data of financial time series. The multiplier distribution is found to be independent of the box size s when s is larger than some crossover scale, providing direct evidence of the existence of sca…
We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…
Reduces symplectic Hamiltonian systems to contact systems, realizing Poincaré's dream.
problem Scaling symmetries in Hamiltonian systems.
method Contact reduction of symplectic Hamiltonian systems.
result Generically possible reduction to contact Hamiltonian systems, reducing inputs needed.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
Deep learning models are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on benign inputs. However, under the black-box setting, most existing adversaries often have a poor transferability to attack other defense models. In this work, from the perspective of regarding the advers…
Paper develops heavy-tailed embeddings for better text classification and augmentation.
problem Improving text classification, especially for extreme values.
method Develops heavy-tailed embeddings using multivariate extreme value theory and introduces a scale-invariant classifier.
result The classifier outperforms baselines and generates meaningful augmented text.
We consider inverse curvature flows in hyperbolic space with starshaped initial hypersurface, driven by positive powers of a homogeneous curvature function. The solutions exist for all time and, after rescaling, converge to a sphere.
Generalized algorithm for translation and scale-invariant prediction.
problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.
A new approach to group fairness treats it as a bargaining problem.
problem Fairness in deploying predictors across subpopulations.
method Interpreting fairness as a bargaining problem and proposing relative improvement.
result Relative improvement provides axiomatic justification and finite-sample convergence guarantees.
The study examines stationary surfaces with boundaries and their properties.
problem Investigating stationary surfaces with boundaries and their critical points.
method A generalized bending energy functional is considered, and the first variation is computed. Boundary-value problems are examined, and a characterization of free-boundary surfaces is given.
result Characterization of free-boundary surfaces with rotational symmetry for scaling-invariant functionals.