Study the averaging principle for non-autonomous slow-fast systems and apply it to financial local stochastic volatility models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on slow convergence in geometric variational problems.
Gradient descent slows significantly in over-parameterized single neuron learning.
Gradient method converges locally linearly for overparameterized Gaussian mixtures.
We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit of the slow motion is given by an averaged ordinary differential equation. We then demonstrate that…
This study improves convergence of two-timescale SA under Markovian noise in reinforcement learning.
We consider the stochastic contextual bandit problem with additional regularization. The motivation comes from problems where the policy of the agent must be close to some baseline policy which is known to perform well on the task. To tackle this problem we use a nonparametric model and propose an algorithm splitting t…
In this paper, a metric with G holonomy and slow rate of convergence to the cone metric is constructed on a ball inside the cone over the flag manifold.
In this paper, we study a popular method for inference of the Bradley-Terry model parameters, namely the MM algorithm, for maximum likelihood estimation and maximum a posteriori probability estimation. This class of models includes the Bradley-Terry model of paired comparisons, the Rao-Kupper model of paired comparison…
The paper analyzes why Gaussianization slows down with higher dimensions and proposes a solution.
Gradient descent converges to perfect classification in neural nets for non-separable data.
Online learning algorithms have impressive convergence properties when it comes to risk minimization and convex games on very large problems. However, they are inherently sequential in their design which prevents them from taking advantage of modern multi-core architectures. In this paper we prove that online learning …
We study a class of weakly identifiable location-scale mixture models for which the maximum likelihood estimates based on i.i.d. samples are known to have lower accuracy than the classical error. We investigate whether the Expectation-Maximization (EM) algorithm also converges slowly for these m…
The study uses the Merton model to estimate PD and finds a phase transition affecting convergence speed.
The continuous dynamical system approach to deep learning is explored in order to devise alternative frameworks for training algorithms. Training is recast as a control problem and this allows us to formulate necessary optimality conditions in continuous time using the Pontryagin's maximum principle (PMP). A modificati…
Federated learning enables a large amount of edge computing devices to jointly learn a model without data sharing. As a leading algorithm in this setting, Federated Averaging (\texttt{FedAvg}) runs Stochastic Gradient Descent (SGD) in parallel on a small subset of the total devices and averages the sequences only once …
Latent class model (LCM), which is a finite mixture of different categorical distributions, is one of the most widely used models in statistics and machine learning fields. Because of its non-continuous nature and the flexibility in shape, researchers in practice areas such as marketing and social sciences also frequen…
Dealing with the shear size and complexity of today's massive data sets requires computational platforms that can analyze data in a parallelized and distributed fashion. A major bottleneck that arises in such modern distributed computing environments is that some of the worker nodes may run slow. These nodes a.k.a.~str…
We consider the action of a pseudo-Anosov mapping class on . This action has north-south dynamics and so, under iteration, laminations converge exponentially to the stable lamination. We study the rate of this convergence and give examples of families of pseudo-Anosov mapping classes where the rate go…
The paper challenges the use of decision trees for pointwise inference due to slow convergence rates.
We describe and analyze a simple algorithm for principal component analysis and singular value decomposition, VR-PCA, which uses computationally cheap stochastic iterations, yet converges exponentially fast to the optimal solution. In contrast, existing algorithms suffer either from slow convergence, or computationally…
PrecGD restores linear convergence in over-parameterized nonconvex matrix factorization.
Study shows how anisotropic data affects learning dynamics in phase retrieval.
SFPO optimizes LLM reasoning by repositioning before updating, improving stability and efficiency.
Regularization leads to balancedness in deep linear networks.
Distributed optimization is essential for training large models on large datasets. Multiple approaches have been proposed to reduce the communication overhead in distributed training, such as synchronizing only after performing multiple local SGD steps, and decentralized methods (e.g., using gossip algorithms) to decou…
New method accelerates convergence for entropy-regularized reinforcement learning problems.
Multirate training speeds up neural network fine-tuning.
In this paper, we focus on approaches to parallelizing stochastic gradient descent (SGD) wherein data is farmed out to a set of workers, the results of which, after a number of updates, are then combined at a central master node. Although such synchronized SGD approaches parallelize well in idealized computing environm…
New method improves convergence of spatial filters in neural networks.
Direct proof shows adaptive gradient descent converges near-linearly for convex functions.
Gradient flow with weight decay shows grokking effect in deep learning.
Optimizes convex functions in finite vs infinite dimensions, revealing slow convergence rates.
This work interprets SFA through variational inference, relaxing linearity constraints.
We consider the minimization of a function defined on a Riemannian manifold accessible only through unbiased estimates of its gradients. We develop a geometric framework to transform a sequence of slowly converging iterates generated from stochastic gradient descent (SGD) on to an averaged i…
Improved convergence speed of principal component analysis through modified learning rules.
Many loss functions in representation learning are invariant under a continuous symmetry transformation. For example, the loss function of word embeddings (Mikolov et al., 2013) remains unchanged if we simultaneously rotate all word and context embedding vectors. We show that representation learning models for time ser…
Deep neural networks have gained tremendous popularity in last few years. They have been applied for the task of classification in almost every domain. Despite the success, deep networks can be incredibly slow to train for even moderate sized models on sufficiently large datasets. Additionally, these networks require l…
We study the contraction of a convex immersed plane curve with speed (1/α)k^{α}, where αin(0,1] is a constant and show that, if the blow-up rate of the curvature is of type one, it will converge to a homothetic self-similar solution. We also discuss a special symmetric case of type two blow-up and show that it converge…
Anytime MiniBatch speeds up online distributed optimization by handling slow nodes.
Regularized LAEs learn principal components efficiently.
Study homogenizes equations on parallelizable manifolds using tensor localization and periodicity.
Accelerates optimal transport computation by 10x with spectral insights.
Machine learning, especially deep neural networks, has been rapidly developed in fields including computer vision, speech recognition and reinforcement learning. Although Mini-batch SGD is one of the most popular stochastic optimization methods in training deep networks, it shows a slow convergence rate due to the larg…
During this last decades, several attempts to construct slow invariant manifold of the Lorenz-Krishnamurthy five-mode model of slow-fast interactions in the atmosphere have been made by various authors. Unfortunately, as in the case of many two-time scales singularly perturbed dynamical systems the various asymptotic p…
The paper optimizes portfolios using MACD signals derived from price history.
We develop a 2D travel time tomography method which regularizes the inversion by modeling groups of slowness pixels from discrete slowness maps, called patches, as sparse linear combinations of atoms from a dictionary. We propose to use dictionary learning during the inversion to adapt dictionaries to specific slowness…
Consider a network of agents connected by communication links, where each agent holds a real value. The gossip problem consists in estimating the average of the values diffused in the network in a distributed manner. We develop a method solving the gossip problem that depends only on the spectral dimension of the netwo…