Method tackles missing covariates in large-scale datasets.
problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification. New decompositions misattribute differences between populations, even when outcomes are identical.
problem Misattribution of differences between populations using common functional decompositions.
method Extending the Kitagawa-Oaxaca-Blinder decomposition to nonlinear functional decompositions.
result Functional ANOVA and Accumulated Local Effects can misattribute differences even when outcomes are identical in two populations.
AICov integrates population covariates for better COVID-19 forecasting.
problem Forecasting COVID-19 with broader social context.
method Integrative deep learning framework with LSTM and multiple data sources.
result Improved prediction of COVID-19 cases and deaths with population risk factors.
Bayesian networks learn sub-population differences from data.
problem Inference from a single network structure can be misleading when data populations are heterogeneous.
method A mixture of Bayesian networks where component probabilities depend on individual characteristics.
result Identifies both network structures and demographic predictors of sub-population membership.
Kernel measures similarity of nonlinear causal structures in heterogeneous populations.
problem Learning causal structure in populations with diverse underlying structures.
method Distance covariance-based kernel for measuring similarity of causal structures.
result Kernel enables clustering of homogeneous subpopulations for causal structure learning.
New findings on optimization landscape of Toeplitz covariance estimation.
problem Understanding the geometry of the Gaussian maximum-likelihood objective for Toeplitz covariance estimation.
method Overparameterized Carathéodory representation of positive definite Toeplitz covariance matrices, focusing on both amplitudes and frequencies.
result Joint optimization of amplitudes and frequencies leads to a benign population landscape, allowing for global recovery of the true Toeplitz covariance.
Develops a weighting framework to generalize ITRs from source to target populations.
problem Challenges in generalizing ITRs from a source population to a target population with differing characteristics.
method A robust sample weighting framework using a reproducing kernel Hilbert space to balance covariates and improve ITR learning methods.
result Improves ITR estimation for the target population compared to other weighting methods.
Study connects covariance cleaning theory to information theory for heavy-tailed distributions.
problem Optimizing covariance matrices for heavy-tailed distributions using information theory.
method Minimizing Frobenius norm and information loss between true and estimated covariance matrices.
result Asymptotic regime of large matrices minimizes information loss for Student's t distributions.
Paper tackles CATE estimation with missing treatment info.
problem Challenges in estimating CATE with missing treatment information.
method Developed MTRNet, a novel CATE estimation algorithm using domain adaptation.
result Improves CATE estimation over state-of-the-art methods.
This paper aims at achieving a simultaneously sparse and low-rank estimator from the semidefinite population covariance matrices. We first benefit from a convex optimization which develops l1-norm penalty to encourage the sparsity and nuclear norm to favor the low-rank property. For the proposed estimator, we then p…
New method improves covariance estimation for weighted samples.
problem Improving covariance estimation for weighted sample data.
method Asymptotic non-linear shrinkage formulas for covariance and precision matrix estimators of weighted sample covariances.
result Asymptotic non-linear shrinkage formulas for covariance and precision matrix estimators of weighted sample covariances.
The covariance matrix of a p-dimensional random variable is a fundamental quantity in data analysis. Given n i.i.d. observations, it is typically estimated by the sample covariance matrix, at a computational cost of O(np2) operations. When n,p are large, this computation may be prohibitively slow. Moreover, …
We introduce a general framework for estimation of inverse covariance, or precision, matrices from heterogeneous populations. The proposed framework uses a Laplacian shrinkage penalty to encourage similarity among estimates from disparate, but related, subpopulations, while allowing for differences among matrices. We p…
In this paper, we provide explicit formulas, in terms of the covariances of sample covariances or sample correlations, for the asymptotic covariances of unrotated factor loading estimates and unique variance estimates. These estimates are extracted from least square, principal, iterative principal component, alpha or i…
Algorithm improves SVM classification in non-Euclidean spaces.
problem Limitations of traditional SVM in non-Euclidean spaces.
method Covariance-adjusted SVM using Cholesky Decomposition.
result Cholesky-SVM outperforms traditional SVM in non-Euclidean spaces.
Method improves treatment effect prediction robust to unknown covariate shifts.
problem Estimating heterogeneous treatment effects for different populations.
method Post-processing CATE T-learners with multi-accurate predictors to handle unknown covariate shifts.
result Improves bias and mean squared error in simulations with covariate shifts.
Improved covariance matrix estimation for multiple classes with limited data.
problem Estimating covariance matrices for multiple classes with scarce data.
method Coupled regularized sample covariance matrix estimator (RSCM) that combines pooled SCM and scaled identity matrix for regularization.
result The coupled RSCM estimators outperform cross-validation in classification tasks with comparable accuracy but faster computation.
While studying response trajectory, often the population of interest may be diverse enough to exist distinct subgroups within it and the longitudinal change in response may not be uniform in these subgroups. That is, the timeslope and/or influence of covariates in longitudinal profile may vary among these different sub…
The only input to attain the portfolio weights of global minimum variance portfolio (GMVP) is the covariance matrix of returns of assets being considered for investment. Since the population covariance matrix is not known, investors use historical data to estimate it. Even though sample covariance matrix is an unbiased…
The paper uses distance covariance to improve fairness in machine learning models.
problem Improving fairness in machine learning models.
method Using conditional and distance covariance statistics to assess independence and add a penalty for fairness.
result The method effectively reduces the fairness gap in machine learning models.
New federated method preserves privacy and estimates treatment effects.
problem Privacy-preserving causal inference for multi-site studies.
method Multiply robust nuisance function estimation, transfer learning.
result Efficient and optimal treatment effect estimation under different scenarios.
Neural network models improve ROC curve evaluation of biomarkers, focusing on age's role in physical activity-mortality association.
problem Improving biomarker evaluation using machine learning for complex relationships.
method Proposes neural network-based covariate-adjusted ROC modeling.
result Age has distinct effects on mortality outcomes when physical activity is measured as total activity time.
Regularization has become a primary tool for developing reliable estimators of the covariance matrix in high-dimensional settings. To curb the curse of dimensionality, numerous methods assume that the population covariance (or inverse covariance) matrix is sparse, while making no particular structural assumptions on th…
New methods resolve conflicting treatment effect estimates in health tech assessments.
problem Conflicting conclusions from different sponsors analyzing the same data.
method Arbitrated indirect treatment comparisons (ArMAIC) targeting a common target population.
result Estimates treatment effects in a common target population, resolving the MAIC paradox.
A new framework for robust risk measurement and portfolio optimization.
problem Uncertainty in mean-covariance space and portfolio optimization challenges.
method Modeling uncertainty with Gelbrich distance and prior structural information, related to optimal transport theory.
result Mean-covariance robust portfolio optimization simplifies to Markowitz model with a regularization term.
Optimal data splitting improves covariance matrix estimation in large datasets.
problem Improving large covariance matrix estimation in high-dimensional settings.
method Focus on holdout method, derive closed-form error expression, connect to eigenvalue variance.
result Optimal train-test split scales as square root of matrix dimension.
Combines trial and observational data to improve policy evaluation.
problem External validity of randomized trial results in target populations.
method Uses covariate data to model trial sampling and certifies policy evaluations.
result Valid trial-based policy evaluations under model miscalibration.
A powerful approach for understanding neural population dynamics is to extract low-dimensional trajectories from population recordings using dimensionality reduction methods. Current approaches for dimensionality reduction on neural data are limited to single population recordings, and can not identify dynamics embedde…
The paper develops statistical inference for gradient flows in optimization.
problem Uncertainty quantification along the entire optimization path.
method Uniform central limit theorem and algorithm-aware covariance estimator.
result Asymptotically valid confidence intervals for target parameter.
New estimator handles covariate shift with closed-form solution and super-efficiency.
problem Handling covariate shift in missing data and causal inference problems.
method Minimum Wasserstein distance estimation framework.
result Closed-form expression and super-efficiency relative to semiparametric efficient estimator.
The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common assumption in this situation is that of prior probability shift, that is, once the la…
Develops a method to estimate personalized treatment regimes from summary statistics.
problem Estimating optimal treatment regimes for a target population when individual-level data is unavailable.
method A weighting framework that tailors a treatment regime for the target population using summary statistics.
result Consistent and asymptotically normal estimator for optimal treatment regimes.
The paper calibrates shrinkage covariance estimators for spectral functionals in high dimensions.
problem Calibrating shrinkage covariance estimators for spectral functionals in high dimensions.
method Derives first-order null laws, distribution-free Davis-Kahan bands, and calibrated tests for spectral functionals under shrinkage.
result Calibrated tests and intervals for spectral functionals are provided, addressing the issue of estimation noise and shrinkage bias.
A large dimensional characterization of robust M-estimators of covariance (or scatter) is provided under the assumption that the dataset comprises independent (essentially Gaussian) legitimate samples as well as arbitrary deterministic samples, referred to as outliers. Building upon recent random matrix advances in the…
TILT improves target domain performance by penalizing an auxiliary component on unlabeled target inputs.
problem Improving performance on target domain under covariate shift.
method TILT uses a novel objective function to decompose the source predictor and penalize an auxiliary component on unlabeled target inputs.
result TILT improves target domain performance over source-only training and other baselines.
We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of covariate vectors along with real-valued responses ("labels"). Otherwise, the formulat…
Proposes a method to correct for covariate shift in meta-analysis of randomized trials.
problem Invalidation of standard IPD meta-analysis due to covariate shift across studies.
method Placebo-anchored transport framework that treats source-trial outcomes as proxy signals and target-trial placebo outcomes as gold labels.
result Yields target-identified effect estimates in connected targets and a principled screen--then--transport procedure in disconnected targets.
Estimates causal effect of managed care plans on NYC Medicaid spending.
problem Generalizing causal estimates to a target population not well-represented by randomized studies.
method Conditional cross-design synthesis estimators combining randomized and observational data.
result Estimates causal effect of managed care plans on health care spending.
Performing statistical inference in high-dimension is an outstanding challenge. A major source of difficulty is the absence of precise information on the distribution of high-dimensional estimators. Here, we consider linear regression in the high-dimensional regime p≫n. In this context, we would like to perform in…
Paper tackles efficient risk estimation under dataset shift conditions.
problem Limited data from target population; auxiliary data available.
method Semiparametric efficiency theory; efficient and multiply robust estimators.
result Developed estimators for various dataset shift conditions.
Trimming helps in conformal prediction when it separates anomaly scores.
problem Effectiveness of trimming in conformal prediction under contamination.
method Analyse fixed-threshold trimming as a replacement of the contaminated calibration law with a retained law.
result Trimming helps when it separates anomaly scores, reducing clean-target coverage to a one-dimensional score-CDF transfer problem.
We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the distribution. The eigenvalues of the covariance of a distribution contain basic informati…
A new QDA classifier for high-dimensional data with spiked covariance.
problem Classifying high-dimensional data with distinct covariance matrices.
method Proposes a novel quadratic classification technique with parameters chosen to maximize the fisher-discriminant ratio.
result The proposed classifier outperforms classical R-QDA and requires lower computational complexity.
We investigate the relationship between the structure of a discrete graphical model and the support of the inverse of a generalized covariance matrix. We show that for certain graph structures, the support of the inverse covariance matrix of indicator variables on the vertices of a graph reflects the conditional indepe…
A new method for steering large agent populations efficiently.
problem Controlling the configuration of a swarm of identical, interacting cooperative agents.
method Mean-Field Schrodinger Bridges with Gaussian Mixture Models.
result A highly efficient parameterization to approximate optimal solutions of the MFSB problem in closed form.
Develops a GP framework for age and year-specific mortality surfaces.
problem Learning the covariance structure of age and year-specific mortality surfaces.
method Genetic programming algorithm to search for the most expressive GP kernel.
result Reveals the presence/absence of cohort effects in different populations.
Study high-dimensional covariance matrix estimators for complex portfolios, improving financial metrics.
problem Estimating covariance matrices in high-dimensional portfolios with nested and one-factor structures.
method Combining random matrix theory, free probability, deterministic equivalents, and two-step covariance estimators.
result Two-step estimators improve financial metrics in complex and one-factor covariance models.
GBMixed boosts mixed models for clustered data, estimating mean and variance flexibly.
problem Flexible estimation of mean and variance components in clustered data.
method Gradient Boosting framework for linear mixed models with likelihood-based gradients.
result GBMixed accurately recovers complex nonlinear fixed effects and covariances.