A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Adaptive weights improve physics-informed neural networks and deep operator networks.
problem Training physics-informed neural networks and deep operator networks can be challenging, leading to unsatisfactory accuracy and efficiency.
method Proposes a pointwise adaptive weighting method that balances the residual decay rate across different training points.
result Our proposed approach of balanced residual decay rates offers advantages including bounded weights, high prediction accuracy, fast convergence rate, low training uncertainty, low computational cost, and ease of hyperparameter tuning.
Regularization in the optimization of deep neural networks is often critical to avoid undesirable over-fitting leading to better generalization of model. One of the most popular regularization algorithms is to impose L-2 penalty on the model parameters resulting in the decay of parameters, called weight-decay, and the …
We prove in this paper the linear stability of the celebrated Schwarzschild family of black holes in general relativity: Solutions to the linearisation of the Einstein vacuum equations around a Schwarzschild metric arising from regular initial data remain globally bounded on the black hole exterior and in fact decay to…
Deep neural networks are typically trained by optimizing a loss function with an SGD variant, in conjunction with a decaying learning rate, until convergence. We show that simple averaging of multiple points along the trajectory of SGD, with a cyclical or constant learning rate, leads to better generalization than conv…
Study on massless Vlasov equation on Reissner-Nordström spacetimes, showing decay rates and non-decay phenomena.
problem Analyzing decay and non-decay rates of solutions to the massless Vlasov equation on Reissner-Nordström spacetimes.
method Quantitative analysis of geodesic flow and comparison to wave equation instability results.
result Exponential decay rates in subextremal cases and polynomial rates in extremal cases, with non-decay of transversal derivatives in extremal cases.
In modern supervised learning, many deep neural networks are able to interpolate the data: the empirical loss can be driven to near zero on all samples simultaneously. In this work, we explicitly exploit this interpolation property for the design of a new optimization algorithm for deep learning, which we term Adaptive…
Minimax optimal convergence rates for classes of stochastic convex optimization problems are well characterized, where the majority of results utilize iterate averaged stochastic gradient descent (SGD) with polynomially decaying step sizes. In contrast, SGD's final iterate behavior has received much less attention desp…
We get decay rate of higher derivatives of nonlinear massless Dirac equations with a kind of "good" spin null form. The method we rely on is similar to that of Li and Zang. However, they only give the decay rate of solution itself to nonlinear massless Dirac system.
In this paper, we prove the linear stability to gravitational and electromagnetic perturbations of the Reissner-Nordström family of charged black holes with small charge. Solutions to the linearized Einstein-Maxwell equations around a Reissner-Nordström solution arising from regular initial data remain globally bounded…
The paper studies harmonic map heat flow stability and decay rates.
problem Analyzing stability and decay rates of harmonic map heat flow solutions.
method Use of homogeneous Besov space B˙p,∞pd(Rd) for small initial data and self-similar decay assumption.
result Decay rates for solutions of the harmonic map flow of the form ∥ablau(t)∥L∞(Rd)≤Ct−21 and self-similar decay under stronger initial conditions.
Learning rate decay (lrDecay) is a \emph{de facto} technique for training modern neural networks. It starts with a large learning rate and then decays it multiple times. It is empirically observed to help both optimization and generalization. Common beliefs in how lrDecay works come from the optimization analysis of (S…
This paper concerns integral varifolds of arbitrary dimension in an open subset of Euclidean space satisfying integrability conditions on their first variation. Firstly, the study of pointwise power decay rates almost everywhere of the quadratic tilt-excess is completed by establishing the precise decay rate for two-di…
We consider any pseudo holomorphic integral 2-cycle in an arbitrary almost complex manifold and perform a blow up analysis at an arbitrary point. Building upon a pseudo algebraic blow up (previously introduced by the author) we prove a geometric rate of decay for the mass ratio towards the limiting density, with an exp…
Residual connections significantly boost the performance of deep neural networks. However, there are few theoretical results that address the influence of residuals on the hypothesis complexity and the generalization ability of deep neural networks. This paper studies the influence of residual connections on the hypoth…
In this paper, we give a new sharp generalization bound of lp-MKL which is a generalized framework of multiple kernel learning (MKL) and imposes lp-mixed-norm regularization instead of l1-mixed-norm regularization. We utilize localization techniques to obtain the sharp learning rate. The bound is characterized by the d…
Stochastic (sub)gradient methods require step size schedule tuning to perform well in practice. Classical tuning strategies decay the step size polynomially and lead to optimal sublinear rates on (strongly) convex problems. An alternative schedule, popular in nonconvex optimization, is called \emph{geometric step decay…
Unified learning-rate scale for CNNs and ResNets, avoiding depth imbalance.
problem Challenges in choosing an appropriate learning rate for deep networks, especially as depth increases.
method Introduces Arithmetic-Mean μP (AM-μP), constraining network-wide average pre-activation second moment to a constant scale, combined with residual-aware He fan-in initialization.
result Demonstrates a −3/2 scaling law for learning rates across depths, enabling zero-shot learning-rate transfer.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2∝d approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.
Momentum is a widely used technique for gradient-based optimizers in deep learning. In this paper, we propose a decaying momentum (\textsc{Demon}) rule. We conduct the first large-scale empirical analysis of momentum decay methods for modern neural network optimization, in addition to the most popular learning rate dec…
Study reveals dynamics of neural networks with normalization, weight decay, and SGD.
problem Understanding the equilibrium condition in Spherical Motion Dynamics (SMD).
method Investigates SMD by exploring the cause of equilibrium condition, introducing assumptions, proposing angular update, and verifying theoretical results.
result Proves weight norm and angular update can converge at linear rate under given assumptions.
We extend Eardley and Moncrief's L∞ estimates for the conformally invariant Yang-Mills-Higgs equations to the Einstein cylinder. Our method is to first work on Minkowski space and localise their estimates, and then carry them to the Einstein cylinder by a conformal transformation. By patching local estimates to…
We prove that the Ricci flow that contracts a hyperbolic cusp has curvature decay like one over time squared. In order to do this, we prove a new Li-Yau type differential Harnack inequality for Ricci flow on surfaces.