Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

108217325433 · Jun 202019922001200920172026
48 results for large catapults

The paper explains spikes in training loss as catapults, improving feature learning and generalization.

problem Understanding and improving the training process of neural networks.
method Analysis of training loss spikes in SGD, empirical evidence of catapults, and demonstration of AGOP alignment.
result Catapults in training loss promote feature learning and better generalization.

Study shows deep linear networks can converge to flatter minima at large learning rates.

problem Understanding the implicit bias of deep linear networks at large learning rates.
method Characterization of deep linear networks for binary classification using logistic loss in the large learning rate regime.
result Gradient descent iterates converge to a flatter minimum in the catapult phase for certain data separation conditions.

Large learning rates lead to various implicit biases in nonconvex optimization.

problem Understanding the conditions under which large learning rates yield edge of stability, balancing, and catapult phenomena.
method Developed a global convergence theory for nonconvex functions without globally Lipschitz continuous gradient, focusing on functions with good regularity.
result These implicit biases are more likely to occur in functions with good regularity, and large learning rates favor flatter regions.

Gradient descent dynamics in quadratic regression models are analyzed, revealing five phases: monotonic, catapult, periodic, chaotic, and divergent.

problem Analyzing the dynamics of gradient descent in quadratic regression models.
method Fine-grained bifurcation analysis of gradient descent dynamics using a cubic map parameterized by the step-size.
result Gradient descent dynamics in quadratic regression models exhibit five distinct phases: monotonic, catapult, periodic, chaotic, and divergent.

New insights into how large learning rates affect transformer training dynamics.

problem Understanding how large learning rates impact the training of transformer models.
method Analyzing a simplified linear transformer model with a two-factor product map.
result Large learning rates can lead to various training outcomes including cycles, chaos, or divergence.

Breiman discusses two statistical cultures, advocating for more research on 'before' and 'after' the black box.

problem Statistical modeling lacks exploration of processes before and after the 'black box'.
method Analyzes Breiman's visual metaphor of two statistical cultures.
result Promotes the importance of studying the 'before' and 'after' of data transformations.

The paper shows how warming up the learning rate improves deep learning performance.

problem Improving deep learning performance through better handling of larger learning rates.
method Systematic experiments with SGD and Adam showing the benefits of warmup and different regimes of operation.
result Properly choosing ηextinitη_{ ext{init}} can eliminate the need for warmup and improve performance.

Minimal surfaces with negative curvature found in large spheres.

problem Existence of minimal surfaces with negative curvature in large dimensional spheres.
method Applied Song's strategy to closed Riemann surfaces with large automorphism groups, resulting in almost hyperbolic minimal surfaces.
result Existence of closed minimal surfaces with negative induced curvature in any sphere of large dimension.

An extra large metric is a spherical cone metric with all cone angles greater than 2 pi and every closed geodesic longer than 2pi. We show that every two-dimensional extra large metric can be triangulated with vertices at cone points only. The argument implies the same result for Euclidean and hyperbolic cone metrics, …

2005-09-14abs ↗pdf ↗

Study large deviations for hypoelliptic diffusion on sub-Riemannian manifolds.

problem Large deviations for hypoelliptic diffusion measures on sub-Riemannian manifolds.
method Rough path theory and manifold-valued Malliavin calculus.
result Proved a large deviation principle for pinned hypoelliptic diffusion measures.

GD with large init shows incremental learning in matrix factorization.

problem Understanding GD's behavior with large initial values in matrix factorization.
method Signal-to-noise ratio concepts and inductive arguments.
result Uncovering an incremental learning phenomenon in GD with large initialization.

As it is known in the finance risk and macroeconomics literature, risk-sharing in large portfolios may increase the probability of creation of default clusters and of systemic risk. We review recent developments on mathematical and computational tools for the quantification of such phenomena. Limiting analysis such as …

2014-02-21abs ↗pdf ↗

Study shows how large neural networks avoid overfitting through decoupling of feature learning and complexity growth.

problem Understanding inductive bias and generalization in large neural networks.
method Dynamical mean field theory applied to large two-layer networks.
result Training dynamics of large networks exhibit a separation of timescales, decoupling feature learning and overfitting.

We study large deviations and rare default clustering events in a dynamic large heterogeneous portfolio of interconnected components. Defaults come as Poisson events and the default intensities of the different components in the system interact through the empirical default rate and via systematic effects that are comm…

2013-11-03abs ↗pdf ↗

Large learning rates enhance model robustness and compressibility.

problem Achieving robustness and resource-efficiency in machine learning models.
method Identifying and utilizing large learning rates as a facilitator for robustness and compressibility.
result Large learning rates produce desirable representation properties and compare favorably to other methods.

We study the concept of coarse disjointness and large scale nn-to-11 functions. As a byproduct, we obtain an Ostrand-type characterization of asymptotic dimension for coarse structures. It is shown that properties like finite asymptotic dimension, coarse finitism, large scale weak paracompactness, ect. are all invari…

2015-08-12abs ↗pdf ↗

Study large deviations in fractional volatility models with non-Gaussian volatility.

problem Large deviations in fractional volatility models with non-Gaussian volatility.
method Established a small-noise large deviation principle for log-price.
result Logarithmic call price asymptotics for large strikes in a special case.

Training deep neural networks using a large batch size has shown promising results and benefits many real-world applications. However, the optimizer converges slowly at early epochs and there is a gap between large-batch deep learning optimization heuristics and theoretical underpinnings. In this paper, we propose a no…

2020-02-04abs ↗pdf ↗

Paper presents an efficient algorithm for learning minimax risk classifiers with large-scale data.

problem Efficient learning of minimax risk classifiers for large-scale data with multiple classes.
method Combination of constraint and column generation for efficient learning.
result 10x speedup for general large-scale data and 100x speedup with many classes.

A faster method for visualization recommendations on large datasets.

problem Infeasibility of state-of-the-art vis-rec models on large datasets due to high computational time.
method Reinforcement-learning (RL) framework that identifies optimal statistics within a time budget.
result Significantly reduces time-to-visualize with minimal error compared to baseline approaches.

Paper introduces scalable clustering for large datasets with outliers.

problem Lack of scalable algorithms for large datasets with outliers.
method Provable robust clustering algorithm based on loss minimization for Gaussian mixture models.
result Algorithm provides high accuracy with theoretical guarantees and outperforms existing methods.

Let G be a finitely presented group, and let p be a prime. Then G is 'large' (respectively, 'p-large') if some normal subgroup with finite index (respectively, index a power of p) admits a non-abelian free quotient. This paper provides a variety of new methods for detecting whether G is large or p-large. These relate t…

2007-02-20abs ↗pdf ↗

In this paper, we study the asymptotic behaviors of implied volatility of an affine jump-diffusion model. Let log stock price under risk-neutral measure follow an affine jump-diffusion model, we show that an explicit form of moment generating function for log stock price can be obtained by solving a set of ordinary dif…

2020-02-29abs ↗pdf ↗

The problem of hedging and pricing sequences of contingent claims in large financial markets is studied. Connection between asymptotic arbitrage and behavior of the αα~-~quantile price is shown. The large Black-Scholes model is carefully examined.

2015-12-21abs ↗pdf ↗

Large deviations theory applied to policy gradient methods.

problem Understanding convergence of policy gradient methods in reinforcement learning.
method Large deviation rate function and contraction principle from large deviations theory.
result Convergence properties of policy gradient methods can be extended to various policy parametrizations.

Study examines large deviations in random walks on hyperbolic spaces.

problem Large deviations in random walks on Gromov-hyperbolic spaces.
method Established large deviations results for distance and translation length of random walks.
result Deduced a special case of a conjecture regarding spectral radii of random matrix products.