Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

69138207276 · Jun 202019922001200920172026
48 results for scale-invariant architectures

SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.

problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.

Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.

problem Why Adam performs better with β1=β2β_1 = β_2.
method Formalized gradient scale invariance and proved it for Adam with equal β1β_1 and β2β_2.
result Adam becomes gradient scale invariant of first order if and only if β1=β2β_1 = β_2.

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

New CNN architecture improves pediatric image segmentation by homogenizing pose and size.

problem Challenges in segmenting pediatric images due to pose and size heterogeneity.
method Spatial Transformer Network (STN) for pose and scale invariance, combined with UNet for segmentation.
result Improved pediatric segmentation, especially renal tumor delineation, with accelerated processing.

Three training regimes found for scale-invariant neural networks on the sphere.

problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.

Proves uniqueness of Ricci flow with scaling invariant estimates.

problem Proving uniqueness of Ricci flow with scaling invariant curvature bound.
method Solving Ricci-harmonic map heat flow in unbounded curvature background.
result Complete Ricci flow starting from uniformly non-collapsed, non-negatively curved manifold is unique in dimension three.

Critical points of scale-invariant curvature energies in 4D are analytic.

problem Analyzing critical points of curvature energies in 4D manifolds.
method Applying Noether's theorem to identify conservation laws and lower order elliptic system of PDEs, then using integrability by compensation and interpolation theory.
result Critical points of scale-invariant curvature energies in 4D are analytic.

New bounds on self-normalized martingales improve online linear regression performance.

problem Improving regret bounds in online linear regression.
method Characterizing scale-invariant bounds on self-normalized martingales.
result For d=1d=1, O(logT)O(\log T) doubly-uniform regret is possible; for d>1d>1, sublinear doubly-uniform regret is impossible.

We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…

2017-08-23abs ↗pdf ↗

AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.

problem Premature decay of effective step sizes in momentum-based optimizers for scale-invariant weights.
method Proposes SGDP and AdamP to eliminate the radial component at each optimizer step, preserving convergence properties.
result Uniform gains across multiple benchmarks, improving model performance.

Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …

2019-05-23abs ↗pdf ↗

Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.

problem Finding and characterizing minimizers and critical points of scale-invariant tangent-point energies for closed curves.
method Develops convergence and regularity theories based on fractional Sobolev spaces and new energy functionals.
result Minimizing sequences converge to locally critical embeddings in all but finitely many points, and locally critical embeddings are regular.

We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…

2009-06-23abs ↗pdf ↗

Study shows how near crushing singularities, Kasner-like regions can exist.

problem Understanding spatial volume densities near crushing singularities.
method Relates existence of Kasner-like regions to asymptotics of spatial volume densities under scale-invariant curvature bounds.
result Kasner-like regions can exist near crushing singularities under certain curvature conditions.

New learning dynamics achieve fast convergence in games without needing to know utility scales.

problem Fast convergence guarantees in learning games require prior knowledge of utility scales.
method Developed scale-free and scale-invariant learning dynamics using optimistic follow-the-regularized-leader with adaptive learning rates and clipping techniques.
result Achieved fast convergence rates to Nash and correlated equilibria without prior utility scale knowledge.

We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…

1999-06-25abs ↗pdf ↗

ASAM improves deep neural network generalization by adapting sharpness to scale.

problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.

Unified framework for scale-invariant representation learning using MAPCA.

problem Learning invariant representations in data.
method Metric-Aware Principal Component Analysis (MAPCA) based on generalized eigenproblem.
result MAPCA provides a unified geometric language for various self-supervised learning objectives.

The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.

problem Nonlinear dynamics of minimal resistance in radial fields.
method Analysis of two non-equilibrium scenarios: scale-invariant free expansion and incompressible source flow.
result Incompressible flow acts as a structural regularizer, admitting unique, smooth, and strictly concave solutions.

This paper analyzes implicit bias in Deep Linear Discriminant Analysis.

problem The implicit bias of Deep Linear Discriminant Analysis.
method Analyzing gradient flow on a L-layer diagonal linear network.
result Under balanced initialization, the network transforms additive updates into multiplicative updates, conserving the (2/L) quasi-norm.

Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.

problem Proving long-time existence and topological rigidity for manifolds with pinched scale-invariant integral curvature.
method Proves long-time existence of Ricci flow for manifolds with bounded curvature and pinched scale-invariant integral curvature, converging to a flat metric.
result Flow converges to a flat metric, implying topological rigidity of the manifold.

New approach improves classification guarantees by focusing on direction rather than regression risk.

problem Improving classification guarantees in binary classification problems.
method Establishing a geometric distinction between classification and regression, leveraging scale invariance.
result Improved guarantees for classification risk compared to regression risk.

We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…

2019-05-15abs ↗pdf ↗

Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.

problem Modeling non-Gaussian time-series with stationary increments.
method Complex wavelet transform for scale variations, joint correlation matrix for scale dependencies, second wavelet transform for diagonalization, maximum entropy models conditioned by scattering spectra coefficients.
result Scattering spectra of self-similar processes are scale invariant, allowing statistical testing and generation of new time-series.

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we refer to as universal sound separation, and it is unknown how performance on spe…

2019-05-08abs ↗pdf ↗

Proves planarity and convexity for ancient solutions of mean curvature flow.

problem Ancient solutions of mean curvature flow in higher codimension.
method Parabolically scale-invariant variation of planarity estimate, convexity proof for pinched solutions.
result Characterizes certain pinched complete ancient solutions and shrinkers in higher codimension.

The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a R2R^2 term) are applied to explain the personal income distribution. A scale invariance exists if there is not the R2R^2 term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…

2002-03-20abs ↗pdf ↗

Batch Normalization (BN) has become a cornerstone of deep learning across diverse architectures, appearing to help optimization as well as generalization. While the idea makes intuitive sense, theoretical analysis of its effectiveness has been lacking. Here theoretical support is provided for one of its conjectured pro…

2018-12-10abs ↗pdf ↗

We consider the learning of algorithmic tasks by mere observation of input-output pairs. Rather than studying this as a black-box discrete regression problem with no assumption whatsoever on the input-output mapping, we concentrate on tasks that are amenable to the principle of divide and conquer, and study what are it…

2016-11-08abs ↗pdf ↗

We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…

2019-08-18abs ↗pdf ↗

This article investigates stationary surfaces with boundaries, which arise as the critical points of functionals dependent on curvature. Precisely, a generalized "bending energy" functional W\mathcal{W} is considered which involves a Lagrangian that is symmetric in the principal curvatures. The first variation of $\ma…

2019-12-15abs ↗pdf ↗

SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.

problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.

Generalized algorithm for translation and scale-invariant prediction.

problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.

Using the monotonicity formulas of Colding and Minicozzi, we prove that on any complete, non-parabolic Riemannian manifold (M3,g)(M^3, g) with non-negative Ricci curvature, the asymptotic weighted scaling invariant integral of scalar curvature has an explicit bound in form of asymptotic volume ratio.

2019-02-24abs ↗pdf ↗