New optimizer Eve uses examplewise gradients for better second-moment estimates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
JME continually estimates data moments privately and accurately.
New adaptive stepsize method for stochastic approximation converges to target point.
A new method for uncertainty estimation in neural networks using existing optimization steps.
Enhances DP linear regression using public data moments.
PMT uses public data moments to make DP feasible for unbounded data.
Several new estimation methods have been recently proposed for the linear regression model with observation error in the design. Different assumptions on the data generating process have motivated different estimators and analysis. In particular, the literature considered (1) observation errors in the design uniformly …
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
ADOPT optimizes Adam to converge with any β2 without bounded noise.
We study the problem of estimating the mean of a random vector given a sample of independent, identically distributed points. We introduce a new estimator that achieves a purely sub-Gaussian performance under the only condition that the second moment of exists. The estimator is based on a novel concept of a…
Consider a random vector with finite second moments. If its precision matrix is an M-matrix, then all partial correlations are non-negative. If that random vector is additionally Gaussian, the corresponding Markov random field (GMRF) is called attractive. We study estimation of M-matrices taking the role of inverse sec…
New estimator tackles multi-task linear regression with outliers, avoiding eigenvalue lower bounds.
On compact surfaces, a Green-Wasserstein inequality cannot be improved without the sqrt(log n) factor.
The paper tackles matrix completion in ultra-sparse sampling, improving imputation accuracy.
A new memory-efficient Adam variant reduces second moments when feasible.
A new portfolio optimization method using the Sherman-Morrison identity.
In this paper we introduce an efficient fat-tail measurement framework that is based on the conditional second moments. We construct a goodness-of-fit statistic that has a direct interpretation and can be used to assess the impact of fat-tails on central data conditional dispersion. Next, we show how to use this framew…
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
Paper proposes ClipSMT algorithm for better ATE estimation.
Paper develops methods for statistical inference in SGD with infinite variance.
Study examines robust regression in high dimensions with heavy-tailed data.
In a recent paper [\textit{M. Cristelli, A. Zaccaria and L. Pietronero, Phys. Rev. E 85, 066108 (2012)}], Cristelli \textit{et al.} analysed relation between skewness and kurtosis for complex dynamical systems and identified two power-law regimes of non-Gaussianity, one of which scales with an exponent of 2 and the oth…
Derives moments of PL networks for robust DNNs.
Iteratively reweighted least squares (IRLS) is a widely-used method in machine learning to estimate the parameters in the generalised linear models. In particular, IRLS for L1 minimisation under the linear model provides a closed-form solution in each step, which is a simple multiplication between the inverse of the we…
New algorithm detects changes in heavy-tailed data streams.
Develops MENT for interpreting and detecting changes in network trajectories.
Uniform stability of a learning algorithm is a classical notion of algorithmic stability introduced to derive high-probability bounds on the generalization error (Bousquet and Elisseeff, 2002). Specifically, for a loss function with range bounded in , the generalization error of a -uniformly stable learning a…
We prove that in metric measure spaces where the entropy functional is K-convex along every Wasserstein geodesic any optimal transport between two absolutely continuous measures with finite second moments lives on a non-branching set of geodesics. As a corollary we obtain that in these spaces there exists only one opti…
Detection of power-law behavior and studies of scaling exponents uncover the characteristics of complexity in many real world phenomena. The complexity of financial markets has always presented challenging issues and provided interesting findings, such as the inverse cubic law in the tails of stock price fluctuation di…
New robust estimators achieve subgaussian bounds using VC-dimension.
New method estimates robust mean in high dimensions with minimized outliers.
We consider the problem of extracting a low-dimensional, linear latent variable structure from high-dimensional random variables. Specifically, we show that under mild conditions and when this structure manifests itself as a linear space that spans the conditional means, it is possible to consistently recover the struc…
Sharp threshold found for aligning Gaussian-weighted graphs.
This paper investigates tradeoffs among optimization errors, statistical rates of convergence and the effect of heavy-tailed errors for high-dimensional robust regression with nonconvex regularization. When the additive errors in linear models have only bounded second moment, we show that iteratively reweighted $\ell_1…
Study detects edge correlation between unlabeled random graphs.
Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
Adaptive gradient methods such as Adam have been shown to be very effective for training deep neural networks (DNNs) by tracking the second moment of gradients to compute the individual learning rates. Differently from existing methods, we make use of the most recent first moment of gradients to compute the individual …
Sophisticated volatility models outperform naive portfolio strategies.
Estimation of the covariance matrix has attracted a lot of attention of the statistical research community over the years, partially due to important applications such as Principal Component Analysis. However, frequently used empirical covariance estimator (and its modifications) is very sensitive to outliers in the da…
New method uses rank-conditioned Horvitz-Thompson estimation for unbiased sample reuse in Plackett-Luce best-of-K objective.
In this paper we introduce a new approach to topic modelling that scales to large datasets by using a compact representation of the data and by leveraging the GPU architecture. In this approach, topics are learned directly from the co-occurrence data of the corpus. In particular, we introduce a novel mixture model whic…
We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic dominance (SSD) relation. This compares the inherent dispersion of random retur…
Suppose that we are given a time series where consecutive samples are believed to come from a probabilistic source, that the source changes from time to time and that the total number of sources is fixed. Our objective is to estimate the distributions of the sources. A standard approach to this problem is to model the …
PLUMAGE improves large model training efficiency and stability.
New method finds closest martingale to Brownian motion.
New method speeds up diffusion models without requiring complex assumptions.
Generative Adversarial Networks improve robust statistics for various distributions.
Proposes variational autoencoder for efficient MMSE estimation.