The paper uses distance covariance to improve fairness in machine learning models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on estimating distances between covariance operators and Gaussian processes.
Proofs Fisher-Rao distance on Gaussian covariance manifold.
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
This study approximates distances between Gaussian processes and covariance operators using RKHS.
This note improves correlation stress tests using geodesic distance.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
Method uses random forest with distance covariance for transfer learning in healthcare.
This work examines the sensitivity of energy distance to mean differences compared to covariance differences.
Optimizes sample reweighting to match laws under covariate shift using Wasserstein distance.
New estimator handles covariate shift with closed-form solution and super-efficiency.
Paper estimates non-causal graphical models using covariance extension and transportation distance.
New method estimates covariance in deep heteroscedastic regression without labels.
Proposes a new Sliced-Wasserstein distance for covariance matrices in M/EEG signals.
Meta learns low-rank covariance factors for better uncertainty estimation.
Bayesian nonparametric models improve OOD detection, especially with complex covariance structures.
Covariance and histogram image descriptors provide an effective way to capture information about images. Both excel when used in combination with special purpose distance metrics. For covariance descriptors these metrics measure the distance along the non-Euclidean Riemannian manifold of symmetric positive definite mat…
Statistical modeling of spatiotemporal phenomena often requires selecting a covariance matrix from a covariance class. Yet standard parametric covariance families can be insufficiently flexible for practical applications, while non-parametric approaches may not easily allow certain kinds of prior knowledge to be incorp…
Extends Mahalanobis distance to Banach spaces for anomaly detection.
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
Proposes a new method for fairness in machine learning with multiple protected attributes.
In this paper, we present a simple non-parametric method for learning the structure of undirected graphs from data that drawn from an underlying unknown distribution. We propose to use Brownian distance covariance to estimate the conditional independences between the random variables and encodes pairwise Markov graph. …
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean discrepancies (MMD), that is, distances between embeddings of distributions to reproduc…
A new distance metric compares probability distributions using kernel covariance operators.
GANs mode collapse solved with Bures distance.
Here, a non-linear analysis method is applied rather than classical one to study projective changes of Finsler metrics. More intuitively, a projectively invariant pseudo-distance is introduced and characterized with respect to the Ricci tensor and its covariant derivatives.
Associating genetic markers with a multidimensional phenotype is an important yet challenging problem. In this work, we establish the equivalence between two popular methods: kernel-machine regression (KMR), and kernel distance covariance (KDC). KMR is a semiparametric regression frameworks that models the covariate ef…
Relying on recent advances in statistical estimation of covariance distances based on random matrix theory, this article proposes an improved covariance and precision matrix estimation for a wide family of metrics. The method is shown to largely outperform the sample covariance matrix estimate and to compete with state…
This study improves estimation of locally stationary functional time series using NW method.
Kernel measures similarity of nonlinear causal structures in heterogeneous populations.
Sufficient dimension reduction (SDR) using distance covariance (DCOV) was recently proposed as an approach to dimension-reduction problems. Compared with other SDR methods, it is model-free without estimating link function and does not require any particular distributions on predictors (see Sheng and Yin, 2013, 2016). …
Magnetoencephalography and electroencephalography (M/EEG) can reveal neuronal dynamics non-invasively in real-time and are therefore appreciated methods in medicine and neuroscience. Recent advances in modeling brain-behavior relationships have highlighted the effectiveness of Riemannian geometry for summarizing the sp…
We develop a novel methodology based on the marriage between the Bhattacharyya distance, a measure of similarity across distributions of random variables, and the Johnson-Lindenstrauss Lemma, a technique for dimension reduction. The resulting technique is a simple yet powerful tool that allows comparisons between data-…
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
Polynomial-time algorithm for estimating covariance in corrupted Gaussian data.
This paper studies geodesics between covariance matrices of different ranks using the Bures-Wasserstein metric.
The correlation length-scale next to the noise variance are the most used hyperparameters for the Gaussian processes. Typically, stationary covariance functions are used, which are only dependent on the distances between input points and thus invariant to the translations in the input space. The optimization of the hyp…
Geometric families of low-rank covariances improve flexibility and tractability in high dimensions.
This article proposes a method to consistently estimate functionals of the eigenvalues of the product of two covariance matrices based on the empirical estimates ($\hat C_a=\frac1{n_a}\sum_{i=1}^{n_a} x_i^{(a)}x_i^{(a){\sf T}…
The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in …
A new framework for robust risk measurement and portfolio optimization.
Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.
Sliced inverse regression is a popular tool for sufficient dimension reduction, which replaces covariates with a minimal set of their linear combinations without loss of information on the conditional distribution of the response given the covariates. The estimated linear combinations include all covariates, making res…
The paper develops tests for comparing means in high dimensions with unknown covariance.
Study convergence and approximations of entropic regularized Wasserstein distances for Gaussian and RKHS measures.
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as establ…
Maximum mean discrepancy (MMD), also called energy distance or N-distance in statistics and Hilbert-Schmidt independence criterion (HSIC), specifically distance covariance in statistics, are among the most popular and successful approaches to quantify the difference and independence of random variables, respectively. T…
Measuring conditional independence is one of the important tasks in statistical inference and is fundamental in causal discovery, feature selection, dimensionality reduction, Bayesian network learning, and others. In this work, we explore the connection between conditional independence measures induced by distances on …