Three training regimes found for scale-invariant neural networks on the sphere.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ASAM improves deep neural network generalization by adapting sharpness to scale.
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
New method trains deep networks robustly without adaptive methods.
It is well known that neural networks with rectified linear units (ReLU) activation functions are positively scale-invariant. Conventional algorithms like stochastic gradient descent optimize the neural networks in the vector space of weights, which is, however, not positively scale-invariant. This mismatch may lead to…
In seeking for sparse and efficient neural network models, many previous works investigated on enforcing L1 or L0 regularizers to encourage weight sparsity during training. The L0 regularizer measures the parameter sparsity directly and is invariant to the scaling of parameter values, but it cannot provide useful gradi…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
Scales attention for long contexts in LLMs.
Neural networks can approximate positive homogeneous functions, especially with multiple hidden layers.
Proves uniqueness of Ricci flow with scaling invariant estimates.
Critical points of scale-invariant curvature energies in 4D are analytic.
New bounds on self-normalized martingales improve online linear regression performance.
We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…
Proves conditions for Willmore surfaces to have finite ends or finite total curvature.
In order to study the application of artificial intelligence (AI) to dental imaging, we applied AI technology to classify a set of panoramic radiographs using (a) a convolutional neural network (CNN) which is a form of an artificial neural network (ANN), (b) representative image cognition algorithms that implement scal…
Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …
Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.
We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…
Study shows how near crushing singularities, Kasner-like regions can exist.
New learning dynamics achieve fast convergence in games without needing to know utility scales.
The study explores how Matrix Product States can represent boolean and continuous functions.
Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.
L-CNNs approximate gauge actions, revealing fixed points with no lattice artifacts.
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
It was empirically confirmed by Keskar et al.\cite{SharpMinima} that flatter minima generalize better. However, for the popular ReLU network, sharp minimum can also generalize well \cite{SharpMinimacan}. The conclusion demonstrates that the existing definitions of flatness fail to account for the complex geometry of Re…
We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…
Bilateral trade relationships in the international level between pairs of countries in the world give rise to the notion of the International Trade Network (ITN). This network has attracted the attention of network researchers as it serves as an excellent example of the weighted networks, the link weight being defined …
Unified framework for scale-invariant representation learning using MAPCA.
The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an -path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…
Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.
New approach improves classification guarantees by focusing on direction rather than regression risk.
We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…
Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.
Proves planarity and convexity for ancient solutions of mean curvature flow.
NOTEARS fails to identify true causal relationships from data.
The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a term) are applied to explain the personal income distribution. A scale invariance exists if there is not the term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…
We propose a new high dimensional semiparametric principal component analysis (PCA) method, named Copula Component Analysis (COCA). The semiparametric model assumes that, after unspecified marginally monotone transformations, the distributions are multivariate Gaussian. COCA improves upon PCA and sparse PCA in three as…
New CNN architecture improves pediatric image segmentation by homogenizing pose and size.
Scale invariance, collective behaviours and structural reorganization are crucial for portfolio management (portfolio composition, hedging, alternative definition of risk, etc.). This lack of any characteristic scale and such elaborated behaviours find their origin in the theory of complex systems. There are several me…
We explore the impact of learning paradigms on training deep neural networks for the Travelling Salesman Problem. We design controlled experiments to train supervised learning (SL) and reinforcement learning (RL) models on fixed graph sizes up to 100 nodes, and evaluate them on variable sized graphs up to 500 nodes. Be…
The statistical properties of the multipliers of the absolute returns are investigated using one-minute high-frequency data of financial time series. The multiplier distribution is found to be independent of the box size when is larger than some crossover scale, providing direct evidence of the existence of sca…
We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…
The condition number predicts efficient information encoding in neural units, aiding model fine-tuning.
This article investigates stationary surfaces with boundaries, which arise as the critical points of functionals dependent on curvature. Precisely, a generalized "bending energy" functional is considered which involves a Lagrangian that is symmetric in the principal curvatures. The first variation of $\ma…
Reduces symplectic Hamiltonian systems to contact systems, realizing Poincaré's dream.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
Deep learning models are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on benign inputs. However, under the black-box setting, most existing adversaries often have a poor transferability to attack other defense models. In this work, from the perspective of regarding the advers…