Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

18365371 · Jun 202019922001200920172026
48 results for momentum compression

A new method reduces communication in distributed learning by 32x.

problem High communication costs in distributed machine learning.
method Distributed compressed SGD with Nesterov's momentum and blockwise compression.
result Achieves 32x reduction in communication cost and converges as fast as full-precision SGD.

GMC uses global momentum for sparse communication in distributed learning.

problem Latency and bandwidth limitations in network communication for large-scale deep learning models.
method GMC utilizes global momentum for sparse communication in DMSGD, improving convergence and accuracy.
result GMC and GMC+ achieve higher test accuracy and faster convergence compared to local momentum methods.

Paper introduces MoTEF for faster decentralized optimization with compressed communication.

problem Efficiency bottleneck in decentralized machine learning applications.
method Integrates communication compression with Momentum Tracking and Error Feedback.
result Significantly outperforms existing methods under arbitrary data heterogeneity.

SQuARM-SGD improves decentralized SGD efficiency with momentum.

problem Efficient decentralized training of large-scale models over networks.
method Fixed local SGD steps with Nesterov's momentum, sparsified and quantized updates, locally computed triggering criterion.
result Convergence rate matches vanilla SGD, momentum improves test performance.

FetchSGD reduces communication in federated learning with sketching.

problem Communication bottlenecks and convergence issues in federated learning.
method FetchSGD uses Count Sketch to compress and merge model updates efficiently.
result FetchSGD achieves high compression rates and good convergence without sparse client participation.

Paper improves communication in distributed optimization, reducing worker-to-server data exchanges.

problem Efficiency in server-to-worker communication in distributed optimization.
method MARINA-P, a novel downlink compression method using correlated compressors; M3, combining MARINA-P with uplink compression.
result MARINA-P achieves provably superior server-to-worker communication complexity with increasing number of workers.

Many popular first-order optimization methods (e.g., Momentum, AdaGrad, Adam) accelerate the convergence rate of deep learning models. However, these algorithms require auxiliary parameters, which cost additional memory proportional to the number of parameters in the model. The problem is becoming more severe as deep l…

2019-02-01abs ↗pdf ↗

FedGLOMO accelerates FL convergence for non-convex functions.

problem Efficiently solving non-convex optimization problems in federated learning with client heterogeneity.
method Combines global and local momentum updates to reduce variance and improve convergence rate.
result Achieves O(ε1.5)\mathcal{O}(ε^{-1.5}) convergence to εε-stationary point, compared to O(ε2)\mathcal{O}(ε^{-2}).

In this paper mechanisms of reversion - momentum transition are considered. Two basic nonlinear mechanisms are highlighted: a slow and fast bifurcation. A slow bifurcation leads to the equilibrium evolution, preceded by stability loss delay of a control parameter. A single order parameter is introduced by Markovian cha…

2015-07-11abs ↗pdf ↗

The study finds that factor momentum is significant only at short lags compared to stock momentum.

problem Investigating the relationship between factor momentum and stock momentum.
method Replicated earlier findings and conducted a spanning test controlling for stock momentum and factor exposure.
result Factor momentum is significant only at short lags after controlling for stock momentum and factor exposure.

We test the price momentum effect in the Korean stock markets under the momentum universe shrinkage to subuniverses of the KOSPI 200. Performance of the momentum strategy is not homogeneous with respect to change of the momentum universe. It is found that some submarkets generate the higher momentum returns than other …

2012-11-28abs ↗pdf ↗

Introduces homotopy momentum sections on multisymplectic manifolds.

problem No specific problem stated; focuses on introducing a new concept.
method Introduces a new concept of homotopy momentum sections on multisymplectic manifolds.
result Shows that a gauged nonlinear sigma model with Wess-Zumino term has homotopy momentum section structure.

Customer momentum is a positive relationship between a firm's returns and past returns of its customers.

problem Understanding the relationship between a firm's returns and its customers' past returns.
method Examined customer momentum using a long-short equally-weighted decile portfolio and Fama-French factor models.
result Customer momentum generates significant monthly returns and is statistically significant.

This paper examines momentum spillover across multiple asset classes using only pricing data.

problem Challenges in studying momentum spillover across diverse asset classes due to lack of common characteristics.
method Utilised a linear and interpretable graph learning model to reveal momentum spillover network.
result Network momentum strategy yields a Sharpe ratio of 1.5 and an annual return of 22%.

The paper analyzes how hyperparameters affect SGD with momentum's convergence rate.

problem The role of hyperparameters in SGD with momentum's convergence rate.
method Theoretical analysis using a hyperparameters-dependent stochastic differential equation (hp-dependent SDE).
result The optimal linear rate of convergence depends on both the learning rate and the momentum coefficient.

New quantum state reconstruction method accelerates convergence.

problem Quantum state reconstruction for larger systems.
method Momentum-Inspired Factored Gradient Descent (MiFGD) combining compressed sensing, non-convex optimization, and acceleration.
result Converges to true density matrix at an accelerated linear rate, provably close to the true matrix.

This paper presents generalized momentum mappings for covariant Hamiltonian field theories. The new momentum mappings arise from a generalization of symplectic geometry to LVYL_VY, the bundle of vertically adapted linear frames over the bundle of field configurations YY. Specifically, the generalized field momentum obs…

2001-11-21abs ↗pdf ↗

We give a detailed discussion about existence and uniqueness of Lu's momentum map. More precisely, we introduce the infinitesimal momentum map, and we study its properties. This allows us to describe the theory of reconstruction of the momentum map from the infinitesimal one. We provide the conditions for the uniquenes…

2012-08-07abs ↗pdf ↗

The paper analyzes how momentum affects convergence in stochastic gradient methods.

problem Lack of clear understanding of momentum's impact on convergence and performance.
method Unified analysis of several popular algorithms using the QHM formulation.
result Provides practical guidelines for setting learning rate and momentum parameters.

DEAM optimizes momentum weights dynamically to improve deep learning model training.

problem Errors in momentum weights propagate errors in optimization algorithms like ADAM.
method DEAM computes adaptive momentum weights based on discriminative angles, reducing hyperparameters and introducing a backtrack term.
result DEAM achieves faster convergence rates in both convex and non-convex deep learning model training.

Adapting momentum from optimization to reinforcement learning.

problem Improving the convergence and stability of reinforcement learning algorithms.
method Introducing Momentum Value Iteration (MoVI) by incorporating an average of consecutive state-action value functions, inspired by the concept of momentum in optimization.
result MoVI improves the convergence and stability of reinforcement learning algorithms, as demonstrated by experiments on Atari games.

Sparse learning speeds up neural network training without sacrificing accuracy.

problem Training deep neural networks efficiently while maintaining performance.
method Sparse momentum algorithm that redistributes and grows weights based on momentum magnitude.
result State-of-the-art sparse performance on various datasets with up to 5.61x faster training.

Momentum speeds up evolutionary processes in machine learning.

problem Accelerating convergence in evolutionary dynamics.
method Combining momentum from machine learning with evolutionary dynamics using information divergences as Lyapunov functions.
result Momentum accelerates convergence of evolutionary dynamics, including the replicator equation and Euclidean gradient descent.

EF21-Muon optimizes deep learning with error feedback, improving efficiency and accuracy.

problem Lack of principled distributed frameworks for non-Euclidean LMO-based optimizers.
method Introduces EF21-Muon, a communication-efficient, non-Euclidean LMO-based optimizer with convergence guarantees.
result First efficient distributed implementation of non-Euclidean LMO-based optimizers, achieving up to 7x communication savings.

Unified model learns from both time-series and cross-sectional momentum features.

problem Separate time-series and cross-sectional momentum strategies do not consider concurrent relationships.
method Spatio-Temporal Momentum strategies using neural networks to combine both types of momentum.
result Simple neural network with single fully connected layer generates trading signals for all assets.

The paper analyzes dynamics of momentum in high dimensions with sparse updates.

problem Theoretical analysis of momentum dynamics in high-dimensional sparse settings.
method Theoretical analysis of two models: least squares with sparse inputs and logistic regression with a rare class.
result Characterization of high-dimensional limits of momentum dynamics and phase structure.

One has not any conventional energy-momentum conservation law in Lagrangian field theory, but relations involving different stress-energy-momentum tensors associated with different connections. It is not obvious how to choose the true energy-momentum tensor. This problem is solved in the framework of the multimomentum …

1995-03-22abs ↗pdf ↗

The paper extends a theorem about momentum maps to singular symplectic spaces.

problem Extending a theorem about momentum maps to singular symplectic spaces.
method Using integral affine stratification and equivariant locally trivial fibrations, the paper extends the linear variation theorem to singular values of the momentum map.
result Cohomology classes of symplectic forms on reduced spaces vary linearly within strata.

This paper analyzes momentum Q-learning with finite-sample guarantees.

problem Improving Q-learning performance with momentum schemes.
method Proposes MomentumQ algorithm integrating Nesterov and Polyak's momentum schemes, analyzes convergence for function approximations.
result Establishes finite-sample convergence rates for MomentumQ, demonstrating better performance than vanilla Q-learning.

Momentum improves deep learning generalization by stabilizing noise and learning features.

problem Improving generalization in deep learning models.
method Empirical and theoretical analysis of gradient descent with momentum (GD+M) vs. gradient descent (GD) in binary classification tasks.
result GD+M outperforms GD in generalization, especially in datasets with shared features and varying margins.

Generalizes momentum map to Courant algebroid for constrained mechanics.

problem Generalizing momentum map to new geometric structures.
method Generalized momentum section on Lie algebroid to Courant algebroid, constructed cohomological formulations.
result Identified momentum section in constrained Hamiltonian mechanics with Courant algebroid symmetry.

Momentum SGD fails to track nonstationary optima due to drift amplification.

problem Tracking nonstationary optima in stochastic optimization.
method Theoretical analysis of SGD and momentum variants under strong convexity and smoothness.
result Momentum incurs a drift-amplification penalty that diverges as the momentum parameter approaches 1, leading to systematic lag.

This paper concentrates on the time series momentum or contrarian effects in the Chinese stock market. We evaluate the performance of the time series momentum strategy applied to major stock indices in mainland China and explore the relation between the performance of time series momentum strategies and some firm-speci…

2017-02-07abs ↗pdf ↗