Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

70139209278 · Jun 202019922001200920172026
48 results for minimal decay

SGD and weight decay encourage neural networks to learn low-rank weight matrices.

problem The bias of SGD towards low-rank weight matrices in neural networks.
method The study investigates the effect of SGD and weight decay on the rank of weight matrices in neural networks, both theoretically and empirically.
result Training with SGD and weight decay induces a bias towards rank minimization in weight matrices, which becomes more pronounced with smaller batch sizes and stronger weight decay.

We study minimal graphic functions on complete Riemannian manifolds $\Si$ with non-negative Ricci curvature, Euclidean volume growth and quadratic curvature decay. We derive global bounds for the gradients for minimal graphic functions of linear growth only on one side. Then we can obtain a Liouville type theorem with …

2013-10-08abs ↗pdf ↗

Study on Kähler manifolds connects curvature decay with growth of holomorphic functions.

problem Analyzing properties of Kähler manifolds with nonnegative bisectional curvature.
method Established precise relations among minimal degree, volume growth, and scalar curvature decay.
result Unified understanding of Kähler-Ricci flow through polynomial growth holomorphic functions.

Study on scalar curvature decay on non-compact manifolds linked at infinity.

problem Understanding scalar curvature decay on non-compact manifolds with topological linking at infinity.
method Analyzing polynomial decay, developing obstruction theory, using μμ--bubble exhaustions, and index theory.
result Topological linking at infinity forces polynomial decay of scalar curvature on manifolds of weakly bounded geometry.

New method improves smoothness of minimizing currents near singular points.

problem Improving smoothness of minimizing currents near singular points.
method New method to estimate the full singular set of the foliation by minimizers and proof of superlinear decay of closeness.
result Generic smoothness of minimizers improved to n9εnn-9-\varepsilon_n for n11n \geq 11.

New height estimate for area minimizing currents, leading to unique tangent cones and decay properties.

problem Analyzing singularities of area minimizing currents.
method Height estimate, decay estimates, techniques inspired by previous works.
result Locally area minimizing currents have a unique tangent cone at almost every point and decay rapidly to a unique tangent plane at branch points.

We study the asymptotic Dirichlet problem for ff-minimal graphs in Cartan-Hadamard manifolds MM. ff-minimal hypersurfaces are natural generalizations of self-shrinkers which play a crucial role in the study of mean curvature flow. In the first part of this paper, we prove the existence of ff-minimal graphs with pre…

2016-05-06abs ↗pdf ↗

Kernel Density Estimation is a very popular technique of approximating a density function from samples. The accuracy is generally well-understood and depends, roughly speaking, on the kernel decay and local smoothness of the true density. However concrete statements in the literature are often invoked in very specific …

2019-01-02abs ↗pdf ↗

Study on optimal ReLU networks with weight decay for interpolation.

problem Interpolating data with radially symmetric distributions using shallow ReLU networks.
method Weight decay regularization in infinite neuron, infinite data limit; analysis of growth rates.
result Existence and growth rates of unique radially symmetric minimizers with weight decay.

Study on singularities in area-minimizing currents, proving unique tangent cones and rectifiability.

problem Understanding singularities in area-minimizing currents.
method Fine excess decay theorems and almost monotonicity of a frequency function.
result Unique tangent cones and countably (m2)(m-2)-rectifiable singular set.

SHIFT framework identifies subgroups with large ML model performance decay.

problem Large model performance decay in subgroups when deployed.
method Subgroup-scanning Hierarchical Inference Framework (SHIFT) for performance drift.
result SHIFT identifies interpretable subgroups with large performance decay and suggests targeted actions to mitigate it.

In this paper we prove a local removable singularity theorem for certain minimal laminations with isolated singularities in a Riemannian three-manifold. This removable singularity theorem is the key result used in our proof that a complete, embedded minimal surface in R3\mathbb{R}^3 with quadratic decay of curvature ha…

2013-08-29abs ↗pdf ↗

Paper tackles dynamic pricing in a geometrically decaying environment, achieving better occupancy with lower rates.

problem Minimizing expected loss in a dynamically changing environment with decisions dependent on the data distribution.
method Introduces algorithms for information and loss function settings, using repeated decision deployment to allow mixing of the environment.
result Iteration complexity matches first and zero order stochastic gradient methods up to logarithmic factors.

Novel Adam-family method with decoupled weight decay for training neural networks.

problem Training nonsmooth neural networks with weight decay.
method Proposes a novel Adam-family method with decoupled weight decay, establishing convergence properties and demonstrating superior performance.
result Asymptotically approximates SGD and enhances generalization performance.

Model proposes neural network for continuous time dynamics with inductive biases.

problem Training neural networks for small datasets with nonlinear dynamics.
method Inductive biases on decay rates and frequencies using Koopman operator theory.
result Higher forecasting performance with single short training sequence.

We consider a random walk on the mapping class group of a surface of finite type. We assume that the random walk is determined by a probability measure whose support is finite and generates a non-elementary subgroup HH. We further assume that HH is not consisting only of lifts with respect to any one covering. Then w…

2014-08-02abs ↗pdf ↗

WSD schedule improves model training efficiency by adapting learning rates dynamically.

problem Fixed compute budgets limit training efficiency of language models.
method Introduces a WSD schedule that uses a constant learning rate followed by a rapid decay phase.
result WSD schedule generates a non-traditional loss curve with stable and decay phases.

This paper concerns some stability properties of higher dimensional catenoids in $\rr^{n+1}$ with n3n\ge 3. We prove that higher dimensional catenoids have index one. We use δδ-stablity for minimal hypersurfaces and show that the catenoid is 2n\frac 2n-stable and a complete 2n\frac 2n-stable minimal hypersurface is a …

2007-08-24abs ↗pdf ↗

The paper connects neural collapse and low-rank bias in networks with L2 regularization.

problem Understanding the emergence of low-rank bias and neural collapse in L2-regularized networks.
method Unified theoretical framework linking TCV and rank of weight matrices, proving global optimality of DNC1, and establishing a benign landscape property.
result Zero TCV across intermediate layers minimizes representation cost under natural architectural constraints, and DNC1 is globally optimal.

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

SGD's performance improves with critical batch size, minimizing SFO complexity.

problem Optimizing SGD's performance with batch size and learning rate.
method Analysis of SGD using constant and decaying learning rates, focusing on batch size effects.
result SGD with critical batch size minimizes SFO complexity.

We prove structure theorems for complete manifolds satisfying both the Ricci curvature lower bound and the weighted Poincaré inequality. In the process, a sharp decay estimate for the minimal positive Green's function is obtained. This estimate only depends on the weight function of the Poincaré inequality, and yields …

2007-01-24abs ↗pdf ↗

A novel k-means method for MNAR data improves clustering accuracy.

problem Improving k-means clustering for data missing not at random.
method A magnitude-decaying MNAR scenario-based k-means method with size constraints.
result The method reduces bias in estimated cluster centers and improves clustering accuracy.

New method for training deep neural networks with regularization, converging to better generalization.

problem Improving generalization of deep neural networks through explicit regularization.
method Regularizer Mirror Descent (RMD) method, inspired by convergence properties of stochastic mirror descent (SMD).
result RMD converges to a point close to the minimizer of the cost function, leading to better generalization performance.

We study dynamic hedging of counterparty risk for a portfolio of credit derivatives. Our empirically driven credit model consists of interacting default intensities which ramp up and then decay after the occurrence of credit events. Using the Galtchouk-Kunita-Watanabe decomposition of the counterparty risk price paymen…

2017-09-04abs ↗pdf ↗

In this paper we prove two theorems. The first one is a structure result that describes the extrinsic geometry of an embedded surface with constant mean curvature (possibly zero) in a homogeneously regular Riemannian three-manifold, in any small neighborhood of a point of large almost-maximal curvature. We next apply t…

2014-01-08abs ↗pdf ↗

Unique solutions found for wave-like decaying null infinity equations.

problem Wave-like decaying null infinity equations with spherically symmetric Einstein-scalar-field.
method Local and global unique solutions for small initial data.
result Sharp decaying condition for unique solutions.