A new memory-efficient Adam variant reduces second moments when feasible.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PMT uses public data moments to make DP feasible for unbounded data.
Enhances DP linear regression using public data moments.
New optimizer Eve uses examplewise gradients for better second-moment estimates.
New adaptive stepsize method for stochastic approximation converges to target point.
In a recent paper [\textit{M. Cristelli, A. Zaccaria and L. Pietronero, Phys. Rev. E 85, 066108 (2012)}], Cristelli \textit{et al.} analysed relation between skewness and kurtosis for complex dynamical systems and identified two power-law regimes of non-Gaussianity, one of which scales with an exponent of 2 and the oth…
JME continually estimates data moments privately and accurately.
A new method for uncertainty estimation in neural networks using existing optimization steps.
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
Develops MENT for interpreting and detecting changes in network trajectories.
We prove that in metric measure spaces where the entropy functional is K-convex along every Wasserstein geodesic any optimal transport between two absolutely continuous measures with finite second moments lives on a non-branching set of geodesics. As a corollary we obtain that in these spaces there exists only one opti…
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
ADOPT optimizes Adam to converge with any β2 without bounded noise.
In this paper we introduce an efficient fat-tail measurement framework that is based on the conditional second moments. We construct a goodness-of-fit statistic that has a direct interpretation and can be used to assess the impact of fat-tails on central data conditional dispersion. Next, we show how to use this framew…
New estimator tackles multi-task linear regression with outliers, avoiding eigenvalue lower bounds.
Study detects edge correlation between unlabeled random graphs.
On compact surfaces, a Green-Wasserstein inequality cannot be improved without the sqrt(log n) factor.
The paper tackles matrix completion in ultra-sparse sampling, improving imputation accuracy.
Several new estimation methods have been recently proposed for the linear regression model with observation error in the design. Different assumptions on the data generating process have motivated different estimators and analysis. In particular, the literature considered (1) observation errors in the design uniformly …
Consider a random vector with finite second moments. If its precision matrix is an M-matrix, then all partial correlations are non-negative. If that random vector is additionally Gaussian, the corresponding Markov random field (GMRF) is called attractive. We study estimation of M-matrices taking the role of inverse sec…
A new portfolio optimization method using the Sherman-Morrison identity.
In this paper we introduce a new approach to topic modelling that scales to large datasets by using a compact representation of the data and by leveraging the GPU architecture. In this approach, topics are learned directly from the co-occurrence data of the corpus. In particular, we introduce a novel mixture model whic…
We study the problem of estimating the mean of a random vector given a sample of independent, identically distributed points. We introduce a new estimator that achieves a purely sub-Gaussian performance under the only condition that the second moment of exists. The estimator is based on a novel concept of a…
New method finds closest martingale to Brownian motion.
While stochastic gradient descent (SGD) and variants have been surprisingly successful for training deep nets, several aspects of the optimization dynamics and generalization are still not well understood. In this paper, we present new empirical observations and theoretical results on both the optimization dynamics and…
We adapt to an infinite dimensional ambient space E.R. Reifenberg's epiperimetric inequality and a quantitative version of D. Preiss' second moments computations to establish that the set of regular points of an almost mass minimizing rectifiable chain in is dense in its support, whenever the group of …
Paper relaxes symmetry conditions for universal feature selection in noisy data.
Uniform stability of a learning algorithm is a classical notion of algorithmic stability introduced to derive high-probability bounds on the generalization error (Bousquet and Elisseeff, 2002). Specifically, for a loss function with range bounded in , the generalization error of a -uniformly stable learning a…
Detection of power-law behavior and studies of scaling exponents uncover the characteristics of complexity in many real world phenomena. The complexity of financial markets has always presented challenging issues and provided interesting findings, such as the inverse cubic law in the tails of stock price fluctuation di…
Paper develops methods for statistical inference in SGD with infinite variance.
Derives moments of PL networks for robust DNNs.
Study examines robust regression in high dimensions with heavy-tailed data.
Paper proposes ClipSMT algorithm for better ATE estimation.
Study resolvent convergence for random matrices with general covariance profiles.
Iteratively reweighted least squares (IRLS) is a widely-used method in machine learning to estimate the parameters in the generalised linear models. In particular, IRLS for L1 minimisation under the linear model provides a closed-form solution in each step, which is a simple multiplication between the inverse of the we…
New algorithm detects changes in heavy-tailed data streams.
We study the price dynamics of stocks traded in the NASDAQ market by considering the statistical properties of an ensemble of stocks traded simultaneously. For each trading day of our database, we study the ensemble return distribution by extracting its first two central moments. According to previous results obtained …
The aim of this paper is to quantify and manage systemic risk caused by default contagion in the interbank market. We model the market as a random directed network, where the vertices represent financial institutions and the weighted edges monetary exposures between them. Our model captures the strong degree of heterog…
Kim-Milman flow map stable under regular target measures
New algorithms avoid a dominant lower-order term in heavy-tailed loss settings.
The paper addresses score-mismatched diffusion models and zero-shot conditional samplers.
X in R^D has mean zero and finite second moments. We show that there is a precise sense in which almost all linear projections of X into R^d (for d < D) look like a scale-mixture of spherical Gaussians -- specifically, a mixture of distributions N(0, sigma^2 I_d) where the weight of the particular sigma component is P …
We prove quantitative convergence rates at which discrete Langevin-like processes converge to the invariant distribution of a related stochastic differential equation. We study the setup where the additive noise can be non-Gaussian and state-dependent and the potential function can be non-convex. We show that the key p…
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
The current interpretation of stochastic gradient descent (SGD) as a stochastic process lacks generality in that its numerical scheme restricts continuous-time dynamics as well as the loss function and the distribution of gradient noise. We introduce a simplified scheme with milder conditions that flexibly interprets S…
Representing examples in a way that is compatible with the underlying classifier can greatly enhance the performance of a learning system. In this paper we investigate scalable techniques for inducing discriminative features by taking advantage of simple second order structure in the data. We focus on multiclass classi…
The paper proposes a method for constructing confidence sets that adapt to the cardinality of the smallest component of a mean vector.
Improved prediction algorithm for 'easy' sequences with reduced regret.