New decompositions misattribute differences between populations, even when outcomes are identical.
problem Misattribution of differences between populations using common functional decompositions.
method Extending the Kitagawa-Oaxaca-Blinder decomposition to nonlinear functional decompositions.
result Functional ANOVA and Accumulated Local Effects can misattribute differences even when outcomes are identical in two populations.
Homotopy equivalence between formalities with different covariant derivatives.
problem Formality of Dolgushev depends on covariant derivative choice.
method Proved homotopy equivalence of L∞-morphisms twisted by gauge equivalent elements. result Globalized formalities with different covariant derivatives are homotopic.
This work examines the sensitivity of energy distance to mean differences compared to covariance differences.
problem The sensitivity of energy distance to mean differences compared to covariance differences when distributions are close.
method Analyzes the energy distance in the case where distributions are close, focusing on sensitivity to mean and covariance differences.
result Energy distance is more sensitive to mean differences than covariance differences when distributions are close.
We study covariance matrix estimation for the case of partially observed random vectors, where different samples contain different subsets of vector coordinates. Each observation is the product of the variable of interest with a 0−1 Bernoulli random variable. We analyze an unbiased covariance estimator under this mod…
Covariate shift relaxes the widely-employed independent and identically distributed (IID) assumption by allowing different training and testing input distributions. Unfortunately, common methods for addressing covariate shift by trying to remove the bias between training and testing distributions using importance weigh…
In the covariate shift learning scenario, the training and test covariate distributions differ, so that a predictor's average loss over the training and test distributions also differ. In this work, we explore the potential of extreme dimension reduction, i.e. to very low dimensions, in improving the performance of imp…
Proofs high-dimensional spectrum convergence of weighted sample covariance.
problem High-dimensional spectrum convergence of weighted sample covariance.
method Proposes a new, concise proof with stronger assumptions.
result Spectrum convergence proven for different weight distributions.
The paper proposes a test to assess rater accuracy while accounting for rater covariates.
problem Assessing the accuracy of raters in medical imaging and forensic studies.
method Covariate-adjusted homogeneity test to determine differences in accuracy among multiple rater groups.
result The proposed test identifies statistically significant differences among five participant groups in a face recognition study.
Paper addresses off-policy evaluation and learning with covariate shift.
problem Evaluating and training a new policy using historical data with a covariate shift.
method Derives efficiency bounds and proposes doubly robust estimators for OPE and OPL under covariate shift.
result Proposes estimators for off-policy evaluation and learning under covariate shift.
A method for rank verification in multivariate Gaussian data, improving on existing approaches.
problem Determining the top K means in multivariate Gaussian data with any covariance structure. method Selective inference tools to generalize the two-sided difference-of-means test for any K and covariance structure. result The method provides a generalization for rank verification in multivariate Gaussian data with any covariance structure.
In this proceeding we give an overview of the idea of covariance (or equivariance) featured in the recent development of convolutional neural networks (CNNs). We study the similarities and differences between the use of covariance in theoretical physics and in the CNN context. Additionally, we demonstrate that the simp…
CovRegRF estimates covariance matrix from covariates using random forests.
problem Estimating conditional covariances or correlations among multivariate responses.
method Random forest trees with a custom splitting rule to maximize covariance difference.
result Accurate covariance matrix estimates and controlled Type-1 error.
We consider the problem of joint estimation of structured inverse covariance matrices. We perform the estimation using groups of measurements with different covariances of the same unknown structure. Assuming the inverse covariances to span a low dimensional linear subspace in the space of symmetric matrices, our aim i…
The paper analyzes how re-weighting helps in reducing variance in high-dimensional kernel methods under covariate shifts.
problem The challenge of high-dimensional kernel methods under covariate shifts and the role of re-weighting.
method Derives asymptotic expansion of high-dimensional kernels under covariate shifts, analyzes bias-variance decomposition, and characterizes the regularized kernel.
result Re-weighting helps in decreasing variance and can be seen as a data-dependent regularization.
New hierarchical model improves on standard practice for high-dimensional data.
problem Poor statistical performance in high-dimensional hierarchical models.
method Model effects as exchangeable across covariates and correlated across datasets.
result Empirical Bayes estimator outperforms classic approach in high-dimensional settings.
We propose a novel estimation approach for the covariance matrix based on the l1-regularized approximate factor model. Our sparse approximate factor (SAF) covariance estimator allows for the existence of weak factors and hence relaxes the pervasiveness assumption generally adopted for the standard approximate factor…
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
Choosing a reference group in Oaxaca-Blinder decomposition can reverse conclusions.
problem The choice of reference group in Oaxaca-Blinder decomposition can lead to different conclusions.
method The study uses the Oaxaca-Blinder decomposition to investigate how the choice of reference group affects the results.
result The Oaxaca-Blinder decomposition can yield different conclusions based on the choice of reference group.
FIRE method improves model performance in federated learning by penalizing fragmentation-induced covariate shifts.
problem Performance degradation in federated learning due to data fragmentation and covariate shift.
method FIRE method accumulates fragmentation-induced covariate shift divergences via approximate Fisher information and uses it as a per-fragment loss penalty.
result FIRE outperforms importance weighting and federated learning benchmarks by up to 5.3% on shifted validation sets.
New method groups similar functional covariates for better modeling.
problem Analyzing functional covariates with similar shapes.
method Coefficient shape alignment regularization approach.
result True grouping structure can be accurately identified under certain conditions.
Estimates the effect of time-varying treatments using machine learning.
problem Estimating the impact of time-varying treatments over multiple periods.
method Difference-in-Differences framework with double/debiased machine learning.
result Higher vaccination rates reduce COVID-19 mortality after several weeks.
New similarity measure for covariate shift improves nonparametric regression rates.
problem Improving nonparametric regression under covariate shift.
method Introducing a new similarity measure based on probability ratios.
result Shows a sharper rate of convergence compared to transfer exponent.
The paper proposes AIS for Bayesian inversion of multioutput signals with covariance estimation.
problem Performing uncertainty analysis of covariance matrices in Bayesian inversion problems for multioutput signals.
method Adaptive Importance Sampling (AIS) scheme, split variables, frequentist approach for noise covariance, prior density over covariance matrix.
result Estimation of model parameters and covariance matrix of noise.
Better signal detection in undersampled data using joint and cross covariances.
problem Detecting shared signals in high-dimensional data with limited samples.
method Analysis of three covariance matrices: individual, cross, and joint.
result Joint and cross covariance matrices detect signals earlier than individual covariances.
This paper considers the problem of estimating multiple related Gaussian graphical models from a p-dimensional dataset consisting of different classes. Our work is based upon the formulation of this problem as group graphical lasso. This paper proposes a novel hybrid covariance thresholding algorithm that can effecti…
Estimates linear model from noisy covariates and instruments using spectral regularization.
problem Estimating a linear model from many noisy covariates and instruments.
method Two-stage least squares with spectral regularization of canonical correlations.
result Upper and lower bounds on estimation error, proving optimality of the method with noisy data.
Optimizes clustering in Gaussian mixtures with varying covariance matrices.
problem Clustering with anisotropic Gaussian mixture models where covariance matrices vary.
method Proposes a computationally feasible hard EM type algorithm.
result Achieves optimal clustering rate with few iterations.
New method tackles MNAR missingness in domain adaptation.
problem Handling missingness in both source and target data.
method Reduces MNAR missingness to imputation problem, leveraging recent MNAR imputation methods.
result Developed a novel domain adaptation procedure for MNAR missingness shift.
The study optimizes investment portfolios using deep learning models for variance-covariance estimation.
problem Estimating an appropriate variance-covariance matrix in Modern Portfolio Theory.
method Employed LSTM-RNN and probabilistic deep learning models (DeepVAR, GPVAR) for multivariate forecasting and portfolio optimization.
result LSTM-RNN models generally yield the best performance in terms of information ratio and annualized returns.
A new GNN architecture called coVariance neural network (VNN) improves stability and transferability of covariance matrix analysis.
problem Stability and transferability issues in covariance matrix analysis.
method Developed coVariance neural network (VNN) that operates on sample covariance matrices.
result VNN is more stable and transferable than PCA-based approaches.
This chapter covers different approaches to policy evaluation for assessing the causal effect of a treatment or intervention on an outcome of interest. As an introduction to causal inference, the discussion starts with the experimental evaluation of a randomized treatment. It then reviews evaluation methods based on se…
This short note reviews so-called Natural Gradient Descent (NGD) for multivariate Gaussians. The Fisher Information Matrix (FIM) is derived for several different parameterizations of Gaussians. Careful attention is paid to the symmetric nature of the covariance matrix when calculating derivatives. We show that there ar…
This paper improves computational efficiency in kernel ridge regression under covariate shift.
problem Covariate shift in nonparametric regression.
method Random projections in RKHS to reduce computational demands.
result Significant computational savings can be achieved without compromising learning performance under covariate shift.
Paper presents a new framework for covariance matrix estimation with geometric insights.
problem Challenges in covariance matrix estimation, especially in finding suitable models and efficient estimation methods.
method General framework for linear restrictions on different transformations of the covariance matrix, including matrix logarithm and its inverse.
result Yields an M-estimator with M-estimation allowing for straightforward asymptotic and finite sample analysis. Paper introduces a novel measure to analyze excess error in classification under covariate shift.
problem Analyzing excess error in classification under covariate shift.
method Utilizes vicinity information to characterize excess error.
result Faster or competitive convergence rates compared to previous techniques.
Improves regression models' performance on covariate shift.
problem Out-of-distribution generalization for regression.
method Spectrally adapting the weights of a pre-trained neural regression model.
result Spectral adaptation improves out-of-distribution performance.
Method estimates multiple related Gaussian distributions using Laplacian regularization.
problem Jointly estimate multiple related zero-mean Gaussian distributions.
method Laplacian regularized stratified model fitting with hyper-parameters to encourage covariance closeness.
result The method performs well, especially in low data regimes, as demonstrated in finance, radar, and weather.
Improved covariance matrix estimation for multiple classes with limited data.
problem Estimating covariance matrices for multiple classes with scarce data.
method Coupled regularized sample covariance matrix estimator (RSCM) that combines pooled SCM and scaled identity matrix for regularization.
result The coupled RSCM estimators outperform cross-validation in classification tasks with comparable accuracy but faster computation.
DRSSS method reduces model training costs by eliminating unsafe samples.
problem Reducing storage and training costs for customized models across different environments.
method Combines DR optimization and SSS for covariate shift.
result Models trained on reduced dataset perform similarly to those on full dataset.
We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired samples; this is a common setup in recent bioinformatics experiments, of which we a…
This paper studies geodesics between covariance matrices of different ranks using the Bures-Wasserstein metric.
problem Geodesics between covariance matrices of varying ranks.
method Analyzes the Bures-Wasserstein distance on covariance matrices, completing previous work on geodesics and providing explicit formulas.
result The set of all minimizing geodesics between two covariance matrices is parametrized by a closed unit ball in R(k−r)imes(l−r). A latent force model is a Gaussian process with a covariance function inspired by a differential operator. Such covariance function is obtained by performing convolution integrals between Green's functions associated to the differential operators, and covariance functions associated to latent functions. In the classica…
The paper addresses the reliability of conformal prediction under covariate shift.
problem Ensuring reliable prediction sets under covariate shift.
method Derives upper bounds on training-conditional coverage.
result Offers PAC guarantees for conformal prediction methods.
New estimator handles covariate shift with closed-form solution and super-efficiency.
problem Handling covariate shift in missing data and causal inference problems.
method Minimum Wasserstein distance estimation framework.
result Closed-form expression and super-efficiency relative to semiparametric efficient estimator.
New method measures treatment effects across different groups.
problem Understanding treatment effects across subgroups while accounting for covariates.
method Proposes BGATE, a new parameter for balanced group average treatment effect.
result Demonstrates usefulness of BGATE in estimating treatment heterogeneity.
ULA estimates covariance of log-concave distributions efficiently.
problem Estimating covariance matrices of log-concave distributions efficiently.
method Unadjusted Langevin algorithm (ULA) for sampling and covariance estimation.
result Sample complexity of single-chain ULA is smaller than that of parallel ULA by a logarithmic factor.
The correlation length-scale next to the noise variance are the most used hyperparameters for the Gaussian processes. Typically, stationary covariance functions are used, which are only dependent on the distances between input points and thus invariant to the translations in the input space. The optimization of the hyp…
New insights into how high-dimensional models handle covariate shifts.
problem Covariate shift in high-dimensional random feature regression.
method Exact high-dimensional asymptotics of random feature regression under covariate shift.
result Overparameterized models exhibit enhanced robustness to covariate shift.