Enhances DP linear regression using public data moments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New adaptive stepsize method for stochastic approximation converges to target point.
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
New optimizer Eve uses examplewise gradients for better second-moment estimates.
A new method for uncertainty estimation in neural networks using existing optimization steps.
A new memory-efficient Adam variant reduces second moments when feasible.
PMT uses public data moments to make DP feasible for unbounded data.
JME continually estimates data moments privately and accurately.
Study detects edge correlation between unlabeled random graphs.
ADOPT optimizes Adam to converge with any β2 without bounded noise.
In a recent paper [\textit{M. Cristelli, A. Zaccaria and L. Pietronero, Phys. Rev. E 85, 066108 (2012)}], Cristelli \textit{et al.} analysed relation between skewness and kurtosis for complex dynamical systems and identified two power-law regimes of non-Gaussianity, one of which scales with an exponent of 2 and the oth…
The paper tackles matrix completion in ultra-sparse sampling, improving imputation accuracy.
New estimator tackles multi-task linear regression with outliers, avoiding eigenvalue lower bounds.
Develops MENT for interpreting and detecting changes in network trajectories.
New method finds closest martingale to Brownian motion.
Several new estimation methods have been recently proposed for the linear regression model with observation error in the design. Different assumptions on the data generating process have motivated different estimators and analysis. In particular, the literature considered (1) observation errors in the design uniformly …
We prove that in metric measure spaces where the entropy functional is K-convex along every Wasserstein geodesic any optimal transport between two absolutely continuous measures with finite second moments lives on a non-branching set of geodesics. As a corollary we obtain that in these spaces there exists only one opti…
A new portfolio optimization method using the Sherman-Morrison identity.
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
In this paper we introduce a new approach to topic modelling that scales to large datasets by using a compact representation of the data and by leveraging the GPU architecture. In this approach, topics are learned directly from the co-occurrence data of the corpus. In particular, we introduce a novel mixture model whic…
In this paper we introduce an efficient fat-tail measurement framework that is based on the conditional second moments. We construct a goodness-of-fit statistic that has a direct interpretation and can be used to assess the impact of fat-tails on central data conditional dispersion. Next, we show how to use this framew…
On compact surfaces, a Green-Wasserstein inequality cannot be improved without the sqrt(log n) factor.
Paper develops methods for statistical inference in SGD with infinite variance.
Consider a random vector with finite second moments. If its precision matrix is an M-matrix, then all partial correlations are non-negative. If that random vector is additionally Gaussian, the corresponding Markov random field (GMRF) is called attractive. We study estimation of M-matrices taking the role of inverse sec…
Iteratively reweighted least squares (IRLS) is a widely-used method in machine learning to estimate the parameters in the generalised linear models. In particular, IRLS for L1 minimisation under the linear model provides a closed-form solution in each step, which is a simple multiplication between the inverse of the we…
Paper proposes ClipSMT algorithm for better ATE estimation.
New algorithm detects changes in heavy-tailed data streams.
We study the problem of estimating the mean of a random vector given a sample of independent, identically distributed points. We introduce a new estimator that achieves a purely sub-Gaussian performance under the only condition that the second moment of exists. The estimator is based on a novel concept of a…
The paper proposes a method for constructing confidence sets that adapt to the cardinality of the smallest component of a mean vector.
While stochastic gradient descent (SGD) and variants have been surprisingly successful for training deep nets, several aspects of the optimization dynamics and generalization are still not well understood. In this paper, we present new empirical observations and theoretical results on both the optimization dynamics and…
We adapt to an infinite dimensional ambient space E.R. Reifenberg's epiperimetric inequality and a quantitative version of D. Preiss' second moments computations to establish that the set of regular points of an almost mass minimizing rectifiable chain in is dense in its support, whenever the group of …
Paper relaxes symmetry conditions for universal feature selection in noisy data.
Uniform stability of a learning algorithm is a classical notion of algorithmic stability introduced to derive high-probability bounds on the generalization error (Bousquet and Elisseeff, 2002). Specifically, for a loss function with range bounded in , the generalization error of a -uniformly stable learning a…
Detection of power-law behavior and studies of scaling exponents uncover the characteristics of complexity in many real world phenomena. The complexity of financial markets has always presented challenging issues and provided interesting findings, such as the inverse cubic law in the tails of stock price fluctuation di…
Adaptive gradient methods such as Adam have been shown to be very effective for training deep neural networks (DNNs) by tracking the second moment of gradients to compute the individual learning rates. Differently from existing methods, we make use of the most recent first moment of gradients to compute the individual …
Derives moments of PL networks for robust DNNs.
Study examines robust regression in high dimensions with heavy-tailed data.
Proposes a new normalization method using convolutional neural networks.
Study resolvent convergence for random matrices with general covariance profiles.
We propose a simple method to learn linear causal cyclic models in the presence of latent variables. The method relies on equilibrium data of the model recorded under a specific kind of interventions ("shift interventions"). The location and strength of these interventions do not have to be known and can be estimated f…
Sharp threshold found for aligning Gaussian-weighted graphs.
We study the price dynamics of stocks traded in the NASDAQ market by considering the statistical properties of an ensemble of stocks traded simultaneously. For each trading day of our database, we study the ensemble return distribution by extracting its first two central moments. According to previous results obtained …
The aim of this paper is to quantify and manage systemic risk caused by default contagion in the interbank market. We model the market as a random directed network, where the vertices represent financial institutions and the weighted edges monetary exposures between them. Our model captures the strong degree of heterog…
Kim-Milman flow map stable under regular target measures
We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic dominance (SSD) relation. This compares the inherent dispersion of random retur…
New method learns SIMs with arbitrary monotone activations without strong distributional assumptions.
New algorithms avoid a dominant lower-order term in heavy-tailed loss settings.
The paper addresses score-mismatched diffusion models and zero-shot conditional samplers.