The paper derives oracle inequalities for estimators with fast and slow rates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A single slow-growing tree matches Random Forest's performance.
Paper introduces slow kill for efficient large-scale variable screening.
We consider the stochastic contextual bandit problem with additional regularization. The motivation comes from problems where the policy of the agent must be close to some baseline policy which is known to perform well on the task. To tackle this problem we use a nonparametric model and propose an algorithm splitting t…
Gradient descent slows significantly in over-parameterized single neuron learning.
EM algorithm converges slowly for weakly identifiable Gaussian mixtures.
In this paper, a metric with G holonomy and slow rate of convergence to the cone metric is constructed on a ball inside the cone over the flag manifold.
In this paper we address the following question: Can we approximately sample from a Bayesian posterior distribution if we are only allowed to touch a small mini-batch of data-items for every sample we generate?. An algorithm based on the Langevin equation with stochastic gradients (SGLD) was previously proposed to solv…
In this paper, we study a popular method for inference of the Bradley-Terry model parameters, namely the MM algorithm, for maximum likelihood estimation and maximum a posteriori probability estimation. This class of models includes the Bradley-Terry model of paired comparisons, the Rao-Kupper model of paired comparison…
We consider in a market model the cooperative emergence of value due to a positive feedback between perception of needs and demand. Here we consider also a negative feedback from production of the traded products, and find that this cooperativity is robust, provided that the production rate is slow. Cooperativity is fo…
Different technological domains have significantly different rates of performance improvement. Prior theory indicates that such differing rates should influence the relative speed of diffusion of the products embodying the different technologies since improvement in performance during the diffusion process increases th…
When applied to training deep neural networks, stochastic gradient descent (SGD) often incurs steady progression phases, interrupted by catastrophic episodes in which loss and gradient norm explode. A possible mitigation of such events is to slow down the learning process. This paper presents a novel approach to contro…
The paper analyzes why Gaussianization slows down with higher dimensions and proposes a solution.
Improved fast rates for decision making with forward-KL regularization in contextual bandits.
Gradient method converges locally linearly for overparameterized Gaussian mixtures.
New ensemble method improves model stability exponentially.
We consider the action of a pseudo-Anosov mapping class on . This action has north-south dynamics and so, under iteration, laminations converge exponentially to the stable lamination. We study the rate of this convergence and give examples of families of pseudo-Anosov mapping classes where the rate go…
Hierarchical pretraining with slow-fast ODEs
Empirical risk minimization (ERM) is a fundamental learning rule for statistical learning problems where the data is generated according to some unknown distribution and returns a hypothesis chosen from a fixed class with small loss . In the parametric setting, depending upon $(\ell…
Theory extends optimal learning rates without realizability assumption.
This article derives prognostic expressions for the evolution of globally aggregated economic wealth, productivity, inflation, technological change, innovation and growth. The approach is to treat civilization as an open, non-equilibrium thermodynamic system that dissipates energy and diffuses matter in order to sustai…
Paper analyzes convergence of FedAvg on non-iid data and provides theoretical guarantees.
Consider the problem of pricing options on forwards in energy markets, when spot prices follow a geometric multi-factor model in which several rates of mean reversion appear. In this paper we investigate the role played by slow mean reversion when pricing and hedging options. In particular, we determine both upper and …
Tian and Yau constructed a complete Ricci-flat Kähler metric on the complement of an ample and smooth anticanonical divisor. We inquire into the behaviour of this metric towards the boundary divisor and prove a slow decay rate of the difference to an appropriate explicitely given referential metric.
Study on slow convergence in geometric variational problems.
The paper tackles long-context linear system identification with improved sample complexity bounds.
The paper challenges the use of decision trees for pointwise inference due to slow convergence rates.
Improved stochastic approximation method reduces residual error.
Efficiently simulates slow dynamics of high-dimensional stochastic systems.
For the problem of high-dimensional sparse linear regression, it is known that an -based estimator can achieve a "fast" rate on the prediction error without any conditions on the design matrix, whereas in absence of restrictive conditions on the design matrix, popular polynomial-time methods only guarante…
Paper analyzes normal approximation for two-timescale stochastic algorithms, revealing interaction between fast and slow timescales.
New bounds show polyhedral surrogates are optimal for generalization.
We consider the minimization of a function defined on a Riemannian manifold accessible only through unbiased estimates of its gradients. We develop a geometric framework to transform a sequence of slowly converging iterates generated from stochastic gradient descent (SGD) on to an averaged i…
The aim of this work is to provide fast and accurate approximation schemes for the Monte-Carlo pricing of derivatives in the Lévy LIBOR model of Eberlein and Özkan (2005). Standard methods can be applied to solve the stochastic differential equations of the successive LIBOR rates but the methods are generally slow. We …
Dealing with the shear size and complexity of today's massive data sets requires computational platforms that can analyze data in a parallelized and distributed fashion. A major bottleneck that arises in such modern distributed computing environments is that some of the worker nodes may run slow. These nodes a.k.a.~str…
Boosting is a learning scheme that combines weak prediction rules to produce a strong composite estimator, with the underlying intuition that one can obtain accurate prediction rules by combining "rough" ones. Although boosting is proved to be consistent and overfitting-resistant, its numerical convergence rate is rela…
PrecGD restores linear convergence in over-parameterized nonconvex matrix factorization.
This work interprets SFA through variational inference, relaxing linearity constraints.
Active learning can't improve over passive in certain settings.
TAEs discover slow modes but mix them with max variance modes.
The paper offers error bounds for quantized dynamical models.
During this last decades, several attempts to construct slow invariant manifold of the Lorenz-Krishnamurthy five-mode model of slow-fast interactions in the atmosphere have been made by various authors. Unfortunately, as in the case of many two-time scales singularly perturbed dynamical systems the various asymptotic p…
We develop a 2D travel time tomography method which regularizes the inversion by modeling groups of slowness pixels from discrete slowness maps, called patches, as sparse linear combinations of atoms from a dictionary. We propose to use dictionary learning during the inversion to adapt dictionaries to specific slowness…
A modular cash-overlay rule for allocating between a fixed growth-defensive risky sleeve and interest-bearing cash.
Method learns dynamics of slow variables from stochastic data.
Optimizes convex functions in finite vs infinite dimensions, revealing slow convergence rates.
Regularization leads to balancedness in deep linear networks.
We propose Power Slow Feature Analysis, a gradient-based method to extract temporally slow features from a high-dimensional input stream that varies on a faster time-scale, as a variant of Slow Feature Analysis (SFA) that allows end-to-end training of arbitrary differentiable architectures and thereby significantly ext…