New method trains deep networks robustly without adaptive methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
New CNN architecture improves pediatric image segmentation by homogenizing pose and size.
Scales attention for long contexts in LLMs.
Three training regimes found for scale-invariant neural networks on the sphere.
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
Proves uniqueness of Ricci flow with scaling invariant estimates.
Critical points of scale-invariant curvature energies in 4D are analytic.
New bounds on self-normalized martingales improve online linear regression performance.
We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …
Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.
We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…
Study shows how near crushing singularities, Kasner-like regions can exist.
New learning dynamics achieve fast convergence in games without needing to know utility scales.
We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…
ASAM improves deep neural network generalization by adapting sharpness to scale.
Unified framework for scale-invariant representation learning using MAPCA.
The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.
This paper analyzes implicit bias in Deep Linear Discriminant Analysis.
Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.
New approach improves classification guarantees by focusing on direction rather than regression risk.
We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…
In seeking for sparse and efficient neural network models, many previous works investigated on enforcing L1 or L0 regularizers to encourage weight sparsity during training. The L0 regularizer measures the parameter sparsity directly and is invariant to the scaling of parameter values, but it cannot provide useful gradi…
Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
It is well known that neural networks with rectified linear units (ReLU) activation functions are positively scale-invariant. Conventional algorithms like stochastic gradient descent optimize the neural networks in the vector space of weights, which is, however, not positively scale-invariant. This mismatch may lead to…
Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we refer to as universal sound separation, and it is unknown how performance on spe…
Proves planarity and convexity for ancient solutions of mean curvature flow.
NOTEARS fails to identify true causal relationships from data.
The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a term) are applied to explain the personal income distribution. A scale invariance exists if there is not the term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…
Batch Normalization (BN) has become a cornerstone of deep learning across diverse architectures, appearing to help optimization as well as generalization. While the idea makes intuitive sense, theoretical analysis of its effectiveness has been lacking. Here theoretical support is provided for one of its conjectured pro…
We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after unspecified marginally monotone transformations, the distributions are multivariate Gaussian. COCA improves upon PCA and sparse PCA in three as…
The statistical properties of the multipliers of the absolute returns are investigated using one-minute high-frequency data of financial time series. The multiplier distribution is found to be independent of the box size when is larger than some crossover scale, providing direct evidence of the existence of sca…
We consider the learning of algorithmic tasks by mere observation of input-output pairs. Rather than studying this as a black-box discrete regression problem with no assumption whatsoever on the input-output mapping, we concentrate on tasks that are amenable to the principle of divide and conquer, and study what are it…
We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…
This article investigates stationary surfaces with boundaries, which arise as the critical points of functionals dependent on curvature. Precisely, a generalized "bending energy" functional is considered which involves a Lagrangian that is symmetric in the principal curvatures. The first variation of $\ma…
Reduces symplectic Hamiltonian systems to contact systems, realizing Poincaré's dream.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
Deep learning models are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on benign inputs. However, under the black-box setting, most existing adversaries often have a poor transferability to attack other defense models. In this work, from the perspective of regarding the advers…
We consider inverse curvature flows in hyperbolic space with starshaped initial hypersurface, driven by positive powers of a homogeneous curvature function. The solutions exist for all time and, after rescaling, converge to a sphere.
Many real-world time series, such as in health, have changepoints where the system's structure or parameters change. Since changepoints can indicate critical events such as onset of illness, it is highly important to detect them. However, existing methods for changepoint detection (CPD) often require user-specified mod…
Generalized algorithm for translation and scale-invariant prediction.
A new approach to group fairness treats it as a bargaining problem.
Using the monotonicity formulas of Colding and Minicozzi, we prove that on any complete, non-parabolic Riemannian manifold with non-negative Ricci curvature, the asymptotic weighted scaling invariant integral of scalar curvature has an explicit bound in form of asymptotic volume ratio.