Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

87173260346 · May 202619922001200920182026
48 results for dependency decay

Analyzes dependencies in sequential datasets to improve deep neural architectures.

problem Improving deep recurrent neural architectures by understanding long distance dependencies.
method Detailed analysis of dependency decay curves in various datasets, testing factors affecting decay, generating synthesized datasets.
result Factors influencing dependency decay curves (number of unique symbols, dataset size, interacting symbols, distance between symbols) can inform optimal hyper-parameters.

Last SGD iterate bounds for overparameterized linear regression.

problem Analyzing the last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression.
method Problem-dependent analysis of last iterate risk bounds of SGD with geometrically decaying stepsize.
result Proved nearly matching upper and lower bounds on the excess risk for last iterate SGD with geometrically decaying stepsize.

Weight decay stabilizes training dynamics by slowing progressive sharpening.

problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.

Regularization and data augmentation can be class-dependent, leading to poor performance on some classes.

problem Class-dependent effects of regularization and data augmentation.
method Evaluation of regularization and data augmentation techniques on Imagenet and INaturalist datasets.
result Regularization and data augmentation can lead to significant performance drops on some classes.

The paper proves scalar curvature decay for uniformly contractible manifolds with finite asymptotic dimension.

problem Proving decay of scalar curvature for uniformly contractible manifolds with finite asymptotic dimension.
method Using index pairing between Dirac operators and compactly supported vector bundles with Lipschitz control, and Lipschitz control for topological K-theory of finite dimensional simplicial complexes.
result The scalar curvature decays to zero at a rate depending only on the contractibility radius and the diameter control of the asymptotic dimension.

Step decay schedules improve convergence in non-convex optimization.

problem Improving convergence in non-convex optimization problems.
method Analyzing convergence rates of step decay schedules in non-convex, convex, and strongly convex problems.
result Step decay schedules achieve O(lnT/T)\mathcal{O}(\ln T/\sqrt{T}) convergence rates in various optimization scenarios.

Paper tackles dynamic pricing in a geometrically decaying environment, achieving better occupancy with lower rates.

problem Minimizing expected loss in a dynamically changing environment with decisions dependent on the data distribution.
method Introduces algorithms for information and loss function settings, using repeated decision deployment to allow mixing of the environment.
result Iteration complexity matches first and zero order stochastic gradient methods up to logarithmic factors.

Optimal multistage method solves noisy minimax problems.

problem Minimizing/maximizing in noisy conditions with smooth and strongly convex-strongly concave settings.
method Multistage Stochastic Gradient Descent Ascent (M-GDA) and Optimistic Gradient Descent Ascent (M-OGDA).
result Achieves optimal linear decay rate with respect to initial error and condition number.

The paper explores how the probability of default estimation changes with temporal correlation decay.

problem Difficulty in estimating the probability of default due to correlations between borrowers.
method Hierarchical Bayesian estimation using beta binomial distribution with temporal correlation.
result A phase transition occurs in the PD estimator, with convergence depending on the power decay index of temporal correlation.

Gradient descent outperforms ridge regression under certain covariance matrix decay conditions.

problem Comparing the performance of gradient descent and ridge regression in linear models.
method Investigated gradient descent and ridge regression for linear regression with random isotropic ground truth.
result Gradient descent outperforms ridge regression under specific covariance matrix decay conditions.

The paper improves transformer generalization bounds using rank-dependent covering number bounds.

problem Improving generalization bounds for transformers.
method Introducing rank-dependent covering number bounds for linear function classes and applying them to transformers.
result Generalization error bounds for transformers decay as O(1/n)O(1/\sqrt{n}) and O(logrw)O(\log r_w), improving existing bounds.

Adaptive time decay functions improve financial product recommendation accuracy.

problem Inaccurate recommendations due to static historical data in finance.
method Time-dependent collaborative filtering with personalized decay functions.
result Significant improvements over state-of-the-art benchmarks in financial product recommendation.

Paper addresses bias in kernel density estimation under minimal assumptions.

problem Kernel density estimation bias under minimal assumptions.
method Demonstrates the need for a balance between kernel decay and bandwidth eigenvalues, and rigorously derives bias bounds.
result Explicit constants and rigorous derivation of bias bounds under minimal assumptions.

The statistical properties of the increments x(t+T) - x(t) of a financial time series depend on the time resolution T on which the increments are considered. A non-parametric approach is used to study the scale dependence of the empirical distribution of the price increments x(t+T) - x(t) of S&P Index futures, for time…

1997-05-08abs ↗pdf ↗

WSD schedule improves model training efficiency by adapting learning rates dynamically.

problem Fixed compute budgets limit training efficiency of language models.
method Introduces a WSD schedule that uses a constant learning rate followed by a rapid decay phase.
result WSD schedule generates a non-traditional loss curve with stable and decay phases.

In this paper we study the behaviour of the continuous spectrum of the Laplacian on a complete Riemannian manifold of bounded curvature under perturbations of the metric. The perturbations that we consider are such that its covariant derivatives up to some order decay with some rate in the geodesic distance from a fixe…

2007-01-10abs ↗pdf ↗

Improved analysis for fair federated learning reduces dependence on noise floor.

problem Asymptotic stationarity in group fair federated learning with reduced noise floor dependence.
method DS FedProxGrad framework with inexact local proximal solutions and fairness regularization.
result Algorithm converges asymptotically to stationarity without dependence on a noise floor.

Regularizers change the geometric properties of loss functions in neural networks.

problem Understanding how different regularizers affect the geometric properties of loss functions in neural networks.
method Examined several regularizers, including weight decay, to determine if the regularized loss function becomes Morse.
result For certain regularizers, the regularized loss function becomes Morse, indicating a change in geometric properties.

Introduces recency bias to improve time-series forecasting.

problem Lack of recency bias in standard Transformer attention for time-series data.
method Reweights attention scores with a smooth heavy-tailed decay to emphasize nearby observations.
result Recency-biased attention consistently improves sequential modeling and achieves competitive performance on time-series forecasting benchmarks.

New approach finds optimal hidden paths in large models, scaling to high dimensions.

problem Finding optimal hidden paths in large, high-dimensional models.
method Developed a new approach to existence of the infinite Viterbi alignment for models satisfying a decay-convexity condition.
result Quantitative bounds on the distance to the infinite Viterbi alignment, demonstrating scalability to high-dimensional problems.

A technique identifies memoryless algorithms approximating memory-dependent optimization methods.

problem Understanding how memory in optimization algorithms affects loss and generalization.
method Introducing a general technique to replace past iterates with the current one and adding a correction term.
result Lion does not have the same implicit anti-regularization as AdamW, explaining its better generalization performance.

We propose the time-dependent generalization of an `ordinary' autonomous human biomechanics, in which total mechanical + biochemical energy is not conserved. We introduce a general framework for time-dependent biomechanics in terms of jet manifolds derived from the extended musculo-skeletal configuration manifold. The …

2009-07-07abs ↗pdf ↗

A new model SMPS alleviates the exponential decay of correlations in MPS.

problem Exponential decay of correlations in Matrix Product States (MPS) limits their power in capturing long-range dependences.
method Introducing long-range interactions (shortcuts) to MPS to decrease correlation length while preserving computational efficiency.
result SMPS can decrease significantly the correlation length of MPS, improving its ability to capture long-range dependences.

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

The paper analyzes multivariate Hawkes processes and their induced population processes.

problem Analyzing the time-dependent joint probability distribution of multivariate Hawkes processes.
method Exact and asymptotic analysis of general multivariate Hawkes processes and their induced population processes.
result Full characterization of the time-dependent joint transform of the multivariate population process and its intensity process.

Using a proprietary dataset of meta-orders and prediction signals, and assuming a quasi-linear impact model, we deconvolve market impact from past correlated trades and a predictable return component to elicit the temporal dependence of the market impact of a single daily meta-order, over a ten day horizon in various e…

2014-07-12abs ↗pdf ↗

Study online learning in RKHS with dependent processes, focusing on \(β\)- and \(φ\)-mixing.

problem Online learning in RKHS with dependent data.
method Online regularized learning algorithm in RKHS, analyzing \(β\)- and \(φ\)-mixing sequences.
result Probabilistic upper bounds and convergence rates for mixing coefficients.

In this article we study the dependence degree of the traded volume of the Dow Jones 30 constituent equities by using a nonextensive generalised form of the Kullback-Leibler information measure. Our results show a slow decay of the dependence degree as a function of the lag. This feature is compatible with the existenc…

2005-10-12abs ↗pdf ↗

HyperAdam learns to optimize neural networks by combining Adam's updates with varying decay rates.

problem Limitation of learned black-box optimizers in generalization ability.
method HyperAdam combines learned and traditional Adam optimizer, adaptively learning weights and decay rates.
result HyperAdam outperforms traditional optimizers in various network training tasks.

Stochastic momentum methods trade compute efficiency for serial runtime.

problem Stochastic momentum methods trade compute efficiency for serial runtime.
method Stochastic HB and ASGD for consistent linear regression with Gaussian covariates.
result HB preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale.

Paper develops an online learning algorithm for functional data models.

problem Recovering slope functions or predictors in functional data models.
method Online regularized learning algorithm in reproducing kernel Hilbert spaces with polynomially decaying step-size.
result Established fast convergence rates for estimation error without capacity assumption.

Weibull weight-scale parameter λλ evolves during AdamW training, with alignment, injection, and decay forces driving its growth and relaxation.

problem Understanding the evolution of the Weibull weight-scale parameter λλ during AdamW training.
method Deriving a leading-order three-force decomposition of the squared weight norm from AdamW updates.
result The alignment force dominates the rise phase, contributing 88-94% of the absolute force budget across four random seeds.

New height estimate for area minimizing currents, leading to unique tangent cones and decay properties.

problem Analyzing singularities of area minimizing currents.
method Height estimate, decay estimates, techniques inspired by previous works.
result Locally area minimizing currents have a unique tangent cone at almost every point and decay rapidly to a unique tangent plane at branch points.