Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

4.5%9.1%13.6%18.2% · Dec 199419922001200920172026
48 results for scale-invariant weights

AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.

problem Premature decay of effective step sizes in momentum-based optimizers for scale-invariant weights.
method Proposes SGDP and AdamP to eliminate the radial component at each optimizer step, preserving convergence properties.
result Uniform gains across multiple benchmarks, improving model performance.

We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…

2017-08-23abs ↗pdf ↗

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

Using the monotonicity formulas of Colding and Minicozzi, we prove that on any complete, non-parabolic Riemannian manifold (M3,g)(M^3, g) with non-negative Ricci curvature, the asymptotic weighted scaling invariant integral of scalar curvature has an explicit bound in form of asymptotic volume ratio.

2019-02-24abs ↗pdf ↗

Three training regimes found for scale-invariant neural networks on the sphere.

problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.

Bilateral trade relationships in the international level between pairs of countries in the world give rise to the notion of the International Trade Network (ITN). This network has attracted the attention of network researchers as it serves as an excellent example of the weighted networks, the link weight being defined …

2007-07-30abs ↗pdf ↗

Proves uniqueness of Ricci flow with scaling invariant estimates.

problem Proving uniqueness of Ricci flow with scaling invariant curvature bound.
method Solving Ricci-harmonic map heat flow in unbounded curvature background.
result Complete Ricci flow starting from uniformly non-collapsed, non-negatively curved manifold is unique in dimension three.

Critical points of scale-invariant curvature energies in 4D are analytic.

problem Analyzing critical points of curvature energies in 4D manifolds.
method Applying Noether's theorem to identify conservation laws and lower order elliptic system of PDEs, then using integrability by compensation and interpolation theory.
result Critical points of scale-invariant curvature energies in 4D are analytic.

New bounds on self-normalized martingales improve online linear regression performance.

problem Improving regret bounds in online linear regression.
method Characterizing scale-invariant bounds on self-normalized martingales.
result For d=1d=1, O(logT)O(\log T) doubly-uniform regret is possible; for d>1d>1, sublinear doubly-uniform regret is impossible.

This paper extends compositional data analysis using graph signal processing.

problem Traditional log-ratios between all variables are not suitable for specific variable relationships.
method Linking compositional data analysis with graph signal processing, it considers only selected log-ratios.
result The approach retains desirable properties of scale invariance and compositional coherence.

Power iteration has been generalized to solve many interesting problems in machine learning and statistics. Despite its striking success, theoretical understanding of when and how such an algorithm enjoys good convergence property is limited. In this work, we introduce a new class of optimization problems called scale …

2019-05-23abs ↗pdf ↗

Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.

problem Finding and characterizing minimizers and critical points of scale-invariant tangent-point energies for closed curves.
method Develops convergence and regularity theories based on fractional Sobolev spaces and new energy functionals.
result Minimizing sequences converge to locally critical embeddings in all but finitely many points, and locally critical embeddings are regular.

We study scale invariant but not necessarily conformal invariant deformations of non-relativistic conformal field theories from the dual gravity viewpoint. We present the corresponding metric that solves the Einstein equation coupled with a massive vector field. We find that, within the class of metric we study, when w…

2009-06-23abs ↗pdf ↗

New learning dynamics achieve fast convergence in games without needing to know utility scales.

problem Fast convergence guarantees in learning games require prior knowledge of utility scales.
method Developed scale-free and scale-invariant learning dynamics using optimistic follow-the-regularized-leader with adaptive learning rates and clipping techniques.
result Achieved fast convergence rates to Nash and correlated equilibria without prior utility scale knowledge.

Study shows how near crushing singularities, Kasner-like regions can exist.

problem Understanding spatial volume densities near crushing singularities.
method Relates existence of Kasner-like regions to asymptotics of spatial volume densities under scale-invariant curvature bounds.
result Kasner-like regions can exist near crushing singularities under certain curvature conditions.

Adam performs better with equal momentum parameters, revealing a gradient scale invariance principle.

problem Why Adam performs better with β1=β2β_1 = β_2.
method Formalized gradient scale invariance and proved it for Adam with equal β1β_1 and β2β_2.
result Adam becomes gradient scale invariant of first order if and only if β1=β2β_1 = β_2.

SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.

problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.

The importance of the power law has been well realized in econophysics over the last decade. For instance, the distribution of the rate of stock price variation and of personal assets show the power law. While these results reveal the striking scale invariance of financial markets, the behaviour of price in real econom…

2006-11-14abs ↗pdf ↗

The condition number predicts efficient information encoding in neural units, aiding model fine-tuning.

problem Efficient information encoding in neural units for various tasks and input modalities.
method Linking the condition number to the log-volume scaling factor and entropy of the output distribution.
result High condition number indicates efficient encoding, reducing overall information transfer.

We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…

1999-06-25abs ↗pdf ↗

ASAM improves deep neural network generalization by adapting sharpness to scale.

problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.

The study proves a new inequality and formula for manifolds with non-negative Ricci curvature.

problem Proving a sharp mean value inequality for non-negative superharmonic functions.
method Develops a new sharp mean value inequality and an explicit formula for weighted scalar curvature.
result The new inequality removes the radius restriction of Schoen-Yau's result and provides an explicit formula for integral of weighted scalar curvature.

It was empirically confirmed by Keskar et al.\cite{SharpMinima} that flatter minima generalize better. However, for the popular ReLU network, sharp minimum can also generalize well \cite{SharpMinimacan}. The conclusion demonstrates that the existing definitions of flatness fail to account for the complex geometry of Re…

2019-03-06abs ↗pdf ↗

Unified framework for scale-invariant representation learning using MAPCA.

problem Learning invariant representations in data.
method Metric-Aware Principal Component Analysis (MAPCA) based on generalized eigenproblem.
result MAPCA provides a unified geometric language for various self-supervised learning objectives.

The paper studies minimal resistance dynamics in radial fields, finding unique solutions for incompressible flows.

problem Nonlinear dynamics of minimal resistance in radial fields.
method Analysis of two non-equilibrium scenarios: scale-invariant free expansion and incompressible source flow.
result Incompressible flow acts as a structural regularizer, admitting unique, smooth, and strictly concave solutions.

Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.

problem Proving long-time existence and topological rigidity for manifolds with pinched scale-invariant integral curvature.
method Proves long-time existence of Ricci flow for manifolds with bounded curvature and pinched scale-invariant integral curvature, converging to a flat metric.
result Flow converges to a flat metric, implying topological rigidity of the manifold.

New approach improves classification guarantees by focusing on direction rather than regression risk.

problem Improving classification guarantees in binary classification problems.
method Establishing a geometric distinction between classification and regression, leveraging scale invariance.
result Improved guarantees for classification risk compared to regression risk.

We develop a regularity theory for extremal knots of scale invariant knot energies defined by J. O'hara in 1991. This class contains as a special case the Möbius energy. For the Möbius energy, due to the celebrated work of Freedman, He, and Wang, we have a relatively good understanding. Their approch is crucially based…

2019-05-15abs ↗pdf ↗

A new metric framework for weighted projective spaces improves clustering and analysis.

problem Proximity measurement in weighted projective spaces with intrinsic scaling and topology.
method Hierarchical clustering framework based on Finsler geometry, quotienting weighted scaling action.
result The constructed metric dFd_F satisfies the triangle inequality, making it a genuine metric.

This paper analyzes implicit bias in Deep Linear Discriminant Analysis.

problem The implicit bias of Deep Linear Discriminant Analysis.
method Analyzing gradient flow on a L-layer diagonal linear network.
result Under balanced initialization, the network transforms additive updates into multiplicative updates, conserving the (2/L) quasi-norm.

Wavelet scattering spectra model non-Gaussian time-series, proving scale invariance for self-similar processes.

problem Modeling non-Gaussian time-series with stationary increments.
method Complex wavelet transform for scale variations, joint correlation matrix for scale dependencies, second wavelet transform for diagonalization, maximum entropy models conditioned by scattering spectra coefficients.
result Scattering spectra of self-similar processes are scale invariant, allowing statistical testing and generation of new time-series.

A new method detects small holes in noisy data.

problem Detecting small holes in high-density regions from noise.
method Robust Density-Aware Distance (RDAD) filtration, incorporating distance-to-measure concept.
result The RDAD filtration prolongs the persistences of small holes, making them distinguishable from noise.

The study improves Poincaré and log-Sobolev inequalities on hyperbolic spaces.

problem Improving Poincaré and log-Sobolev inequalities on hyperbolic spaces.
method Establishing scale-dependent Poincaré-Hardy type identities and choosing suitable parameters, potentials, and vector fields.
result Derives new versions and substantially improves existing inequalities.

Proves planarity and convexity for ancient solutions of mean curvature flow.

problem Ancient solutions of mean curvature flow in higher codimension.
method Parabolically scale-invariant variation of planarity estimate, convexity proof for pinched solutions.
result Characterizes certain pinched complete ancient solutions and shrinkers in higher codimension.

The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a R2R^2 term) are applied to explain the personal income distribution. A scale invariance exists if there is not the R2R^2 term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…

2002-03-20abs ↗pdf ↗

This paper investigates the effectiveness of decoupled weight decay at the start of training.

problem The traditional approach to weight decay is not effective throughout training.
method The authors investigate decoupled weight decay, applying it only at the start of training.
result Applying weight decay only at the start of training stabilizes network weights and improves performance.

We develop an entropic framework to model the dynamics of stocks and European Options. Entropic inference is an inductive inference framework equipped with proper tools to handle situations where incomplete information is available. The objective of the paper is to lay down an alternative framework for modeling dynamic…

2019-08-18abs ↗pdf ↗