Weight decay is one of the standard tricks in the neural network toolbox, but the reasons for its regularization effect are poorly understood, and recent results have cast doubt on the traditional interpretation in terms of L2 regularization. Literal weight decay has been shown to outperform L2 regularization for…
We found that factors decay over time, with momentum fitting best.
problem Understanding how factors decay over time and their impact on performance.
method Derived a hyperbolic decay model for factors, tested against linear and exponential alternatives.
result Momentum exhibits hyperbolic decay, outperforming linear and exponential models.
Study on scalar curvature decay on non-compact manifolds linked at infinity.
problem Understanding scalar curvature decay on non-compact manifolds with topological linking at infinity.
method Analyzing polynomial decay, developing obstruction theory, using μ--bubble exhaustions, and index theory. result Topological linking at infinity forces polynomial decay of scalar curvature on manifolds of weakly bounded geometry.
Unified framework for analyzing gradient flows of measures with exponential decay of entropy.
problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.
This paper reveals periodic behavior in neural network training with BN and weight decay.
problem Understanding the dynamics of neural network training with BN and weight decay.
method Rigorous investigation of empirical and theoretical mechanisms.
result Periodic behavior in training is a generalization of previously opposing perspectives.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
SAD-DPSGD improves model performance on imbalanced medical datasets like HAM10000.
problem Data leakage and imbalanced distribution in medical image classification datasets.
method SAD-DPSGD uses a linear decaying mechanism for noise and clipping thresholds to enhance performance.
result SAD-DPSGD outperforms Auto-DPSGD on HAM10000, improving accuracy by 2.15%.
Analyzes why neural networks generalize beyond training data.
problem Understanding why neural networks generalize beyond training data.
method Examined loss landscapes of neural networks to identify the mismatch between training and test losses.
result Identified the 'LU mechanism' explaining grokking in various tasks.
Gradient flow with weight decay shows grokking effect in deep learning.
problem Understanding the grokking effect in deep learning.
method Analyzing gradient flow dynamics with weight decay.
result Weight decay causes slow norm reduction, explaining grokking.
Paper tackles causal inference with partially labeled data, introducing robust methods.
problem Challenges in causal inference due to partially labeled datasets and potential bias.
method Decaying missing-at-random framework and BRSS estimator for doubly robust causal inference.
result Established asymptotic normality of BRSS estimator under decaying labeling propensity scores.
New model explains volatility after extreme stock market events.
problem Understanding volatility dynamics after extreme stock market events.
method Proposed a new dynamical model using high frequency minute data.
result Volatility after extreme events follows a stretched exponential decay initially and a power law decay later.
Survey on stability of Minkowski spacetime in relativity.
problem Nonlinear stability of Minkowski spacetime in general relativity.
method Decay assumptions, geometric foliations, energy identities, and gauge choices.
result Understanding of decay, dispersion, and geometry-analysis interplay.
We study the problem of what causes prices to change. We define the mechanical impact of a trading order as the change in future prices in the absence of any future changes in decision making, and its it informational impact as the remainder of the total impact once mechanical impact is removed. We introduce a method o…
AdamNX improves Adam's stability by adjusting its learning rate.
problem Adam's tendency to converge to non-flat minima in large-scale models.
method Proposes a novel exponential decay mechanism for Adam's second-order moment estimate.
result AdamNX outperforms Adam and its variants in stability and performance.
Study proposes GRU-D networks for missing value handling in road surface friction prediction.
problem Missing values in road surface friction data affect prediction accuracy.
method Gated Recurrent Unit (GRU) network with decay mechanism.
result GRU-D networks outperform baseline models in road surface friction prediction.
The three-state agent-based 2D model of financial markets as proposed by Giulia Iori has been extended by introducing increasing trust in the correctly predicting agents, a more realistic consultation procedure as well as a formal validation mechanism. This paper shows that such a model correctly reproduces the three f…
Universal model for soft tissue mechanics under shock waves.
problem Modeling shock wave mechanics in soft biological tissues.
method Continuum mixture theory with phase-field mechanics.
result Universal thermodynamically consistent formulation for soft porous tissues.
A new flow method reduces Lorentz contraction to a simple algebraic decay.
problem Reducing Lorentz contraction in geometric models.
method Variational scalar conformal flow with algebraic decay.
result Explicit algebraic decay law for energy functional.
This paper investigates the effectiveness of adversarial training in enhancing the robustness of Deep Q-Network (DQN) policies to state-space perturbations. We first present a formal analysis of adversarial training in DQN agents and its performance with respect to the proportion of adversarial perturbations to nominal…
We show implicit filter level sparsity manifests in convolutional neural networks (CNNs) which employ Batch Normalization and ReLU activation, and are trained with adaptive gradient descent techniques and L2 regularization or weight decay. Through an extensive empirical study (Mehta et al., 2019) we hypothesize the mec…
Introduces recency bias to improve time-series forecasting.
problem Lack of recency bias in standard Transformer attention for time-series data.
method Reweights attention scores with a smooth heavy-tailed decay to emphasize nearby observations.
result Recency-biased attention consistently improves sequential modeling and achieves competitive performance on time-series forecasting benchmarks.
Develops a new deep learning formulation using Mori-Zwanzig formalism.
problem Improves deep learning by introducing a new concept of memory.
method Uses Mori-Zwanzig formalism to propagate quantities of interest through neural networks.
result Rigorously transforms deep networks into shallow ones using decay property of memory operator.
The rich-get-richer mechanism (agents increase their ``wealth'' randomly at a rate proportional to their holdings) is often invoked to explain the Pareto power-law distribution observed in many physical situations, such as the degree distribution of growing scale free nets. We use two different analytical approaches, a…
Stock markets can be characterized by fat tails in the volatility distribution, clustering of volatilities and slow decay of their time correlations. For an explanation models with several mechanisms and consequently many parameters as the Lux-Marchesi model have been used. We show that a simple herding model with only…
In this article we study the dependence degree of the traded volume of the Dow Jones 30 constituent equities by using a nonextensive generalised form of the Kullback-Leibler information measure. Our results show a slow decay of the dependence degree as a function of the lag. This feature is compatible with the existenc…
Logarithmic-time schedules boost large-scale language model training efficiency.
problem Improving performance and efficiency in large-scale language model training.
method Designing time-varying hyperparameters (β1,β2,λ) for AdamW, specifically logarithmic-time scheduling with damping mechanisms. result ADANA optimizer achieves up to 40% compute efficiency compared to tuned AdamW, with gains persisting as model scale increases.
We examine on the static and dynamical properties of quantum knots in a Bose-Einstein condensate. In particular, we consider the Gross-Pitaevskii model and revise a technique to construct ab initio the condensate wave-function of a generic torus knot. After analysing its excitation energy, we study its dynamics relatin…
Transformers simplify modeling of small longitudinal cohort data by reducing parameters and incorporating attention mechanisms.
problem Challenges in modeling longitudinal cohort data due to complex temporal dependencies and large dataset requirements.
method Simplified transformer architecture with attention mechanism, autoregressive model, and kernel-based temporal decay.
result The approach recovers contextual dependencies even with small datasets, identifying temporal patterns in stress and mental health.
Suppose k centers are fit to m points by heuristically minimizing the k-means cost; what is the corresponding fit over the source distribution? This question is resolved here for distributions with p≥4 bounded moments; in particular, the difference between the sample cost and distribution cost decays with $…
WildCat efficiently compresses neural network attention mechanisms.
problem Expensive quadratic runtime of attention mechanisms in neural networks.
method Uses a weighted coreset with a fast subsampling algorithm to approximate attention with near-linear time complexity.
result Approximates exact attention with super-polynomial error decay and near-linear runtime.
This paper proposes a method for modeling event sequences with ambiguous timestamps, a time-discounting convolution. Unlike in ordinary time series, time intervals are not constant, small time-shifts have no significant effect, and inputting timestamps or time durations into a model is not effective. The criteria that …
Study on efficiency of Dutch auctions on blockchains considering various parameters.
problem Efficiency and fairness in Dutch auctions on blockchains.
method Modeling Dutch auctions with Poisson process and geometric Brownian motion, computing expected losses and time-to-fill.
result Tradeoff between speed and quality in Dutch auctions, useful for setting parameters.
We analyze the fluctuations in the gross domestic product (GDP) of 152 countries for the period 1950--1992. We find that (i) the distribution of annual growth rates for countries of a given GDP decays with ``fatter'' tails than for a Gaussian, and (ii) the width of the distribution scales as a power law of GDP with a s…
Framework models supervised learning in non-stationary data.
problem Non-stationary data in supervised learning.
method Statistical physics methods applied to LVQ and neural networks.
result LVQ and ReLU have different sensitivity to concept drift.
We propose the time-dependent generalization of an `ordinary' autonomous human biomechanics, in which total mechanical + biochemical energy is not conserved. We introduce a general framework for time-dependent biomechanics in terms of jet manifolds derived from the extended musculo-skeletal configuration manifold. The …
We propose a stochastic process driven by memory effect with novel distributions including both exponential and leptokurtic heavy-tailed distributions. A class of distribution is analytically derived from the continuum limit of the discrete binary process with the renormalized auto-correlation and the closed form momen…
Extends Minkowski stability proof to minimal decay assumptions.
problem Global stability of Minkowski spacetime with minimal decay.
method Extends Christodoulou-Klainerman's proof to minimal decay assumptions.
result Exterior stability of Minkowski holds with borderline decay.
In this paper, we reveal the attenuation mechanism of anchor of the commodity money from the perspective of logistics warehousing costs, and propose a novel Decayed Commodity Money (DCM) for the store of value across time and space. Considering the logistics cost of commodity warehousing by the third financial institut…
New method creates vacuum data at minimal and borderline decay thresholds.
problem Creating vacuum initial data at specific decay thresholds.
method Conical solution-operator method applied to vacuum asymptotically flat initial data.
result Demonstrates global and exterior stability of Minkowski spacetime.
Unified framework reveals regularization mechanism in deep ReLU networks via convex optimization.
problem Understanding the success of deep neural networks.
method Developed a unified framework using convex optimization to reveal regularization mechanisms.
result ReLU networks can be globally optimized via convex programs, enforcing sparsity.
The study shows that the visible range from a point on harmonic manifolds follows an exponential distribution.
problem Understanding the visible range from a point on harmonic manifolds.
method Analyzing Poisson Boolean models on harmonic manifolds, focusing on the geometric mechanism of tube volumes around geodesic segments.
result The visible range from a point on harmonic manifolds follows an exponential distribution.
High-frequency trading models fail due to overfitting and survivor bias.
problem Failure of hybrid DRL-EC trading systems in high-frequency environments.
method Deployed a population of 500 agents in a high-frequency cryptocurrency environment, analyzing failure modes through multi-disciplinary lens.
result Increasing model complexity without information asymmetry exacerbates systemic fragility.
GradPower speeds up language model training with minimal code changes.
problem Slower training of large language models.
method Elementwise sign-power transformation applied to gradients.
result Consistently lower terminal loss across various models and datasets.
Paper uses AI to predict medications from medical codes, improving accuracy in healthcare.
problem Predicting medications from incomplete or incorrect medical codes is challenging.
method Robust Recurrent Neural Networks (RNNs) with decay mechanism and noise injection.
result The method accurately predicts medication orders from contaminated medical codes.
Unique solutions found for wave-like decaying null infinity equations.
problem Wave-like decaying null infinity equations with spherically symmetric Einstein-scalar-field.
method Local and global unique solutions for small initial data.
result Sharp decaying condition for unique solutions.
New findings on flatness of certain metrics with fast decay.
problem Rigidity of positive mass theorem under fast metric decay.
method Considered metrics with nonnegative scalar curvature and rapid decay at infinity.
result Any such metric is necessarily flat in dimensions 4 and higher if decay rate exceeds Schwarzschild metric.
Develops neural networks for reductive Lie groups, enhancing symmetry respect.
problem Symmetry respect in neural networks for reductive Lie groups.
method General equivariant neural network architecture for any reductive Lie Group G.
result Demonstrates generality and performance in top quark decay tagging and shape recognition.
We propose a stochastic process driven by the memory effect with novel distributions which include both exponential and leptokurtic heavy-tailed distributions. A class of the distributions is analytically derived from the continuum limit of the discrete binary process with the renormalized auto-correlation. The moment …