Paper introduces a novel method for dynamic covariance estimation with random forests.
problem Estimating high-dimensional dynamic covariance matrices with multiple covariates.
method Nonparametric approach using random forests.
result Uniform consistency theory and error rates established for high-dimensional scenarios.
Nonsingular estimation of high dimensional covariance matrices is an important step in many statistical procedures like classification, clustering, variable selection an future extraction. After a review of the essential background material, this paper introduces a technique we call slicing for obtaining a nonsingular …
Proofs high-dimensional spectrum convergence of weighted sample covariance.
problem High-dimensional spectrum convergence of weighted sample covariance.
method Proposes a new, concise proof with stronger assumptions.
result Spectrum convergence proven for different weight distributions.
SPARKLE handles high-dimensional covariates for online decision-making.
problem Complex reward-covariate relationships in high-dimensional settings.
method SPARKLE uses a sparse additive reward model with doubly penalized estimator and adaptive screening.
result SPARKLE achieves sublinear regret bound logarithmic in covariate dimensionality.
Study improves Hayashi-Yoshida estimator for high-dimensional stock covolatility.
problem Inconsistent performance of Hayashi-Yoshida estimator in high dimensions.
method Analyzed the limiting spectral distribution of the Hayashi-Yoshida estimator.
result Established the connection between the estimator's spectrum and the true covariance matrix in high dimensions.
Proposes FarmHazard model for hazard regression with correlated covariates.
problem Model selection challenges in high-dimensional data with correlated covariates.
method Factor-Augmented Regularized Model for Hazard Regression (FarmHazard) that learns latent factors and idiosyncratic components.
result Proves model selection and estimation consistency under mild conditions.
A new QDA classifier for high-dimensional data with spiked covariance.
problem Classifying high-dimensional data with distinct covariance matrices.
method Proposes a novel quadratic classification technique with parameters chosen to maximize the fisher-discriminant ratio.
result The proposed classifier outperforms classical R-QDA and requires lower computational complexity.
New method tackles high-dimensional SBL without covariance matrices.
problem Sparse coding problem in high-dimensional settings.
method Parallel solution of multiple linear systems using conjugate gradient algorithm.
result Our method scales better in computation time and memory.
New insights into how high-dimensional models handle covariate shifts.
problem Covariate shift in high-dimensional random feature regression.
method Exact high-dimensional asymptotics of random feature regression under covariate shift.
result Overparameterized models exhibit enhanced robustness to covariate shift.
The paper analyzes how re-weighting helps in reducing variance in high-dimensional kernel methods under covariate shifts.
problem The challenge of high-dimensional kernel methods under covariate shifts and the role of re-weighting.
method Derives asymptotic expansion of high-dimensional kernels under covariate shifts, analyzes bias-variance decomposition, and characterizes the regularized kernel.
result Re-weighting helps in decreasing variance and can be seen as a data-dependent regularization.
Paper proposes a robust test for high-dimensional models with large covariates and instruments.
problem Testing high-dimensional linear instrumental variable models with large covariates and instruments.
method Introduces a test based on the maximum norm of multiple parameters and a power-enhanced test.
result The proposed test is robust to heteroskedastic errors and has higher power than existing tests.
Proposes spBART for risk prediction using epigenetic signatures and covariates.
problem Complex high-dimensional epigenetic data and low-dimensional covariates for risk prediction.
method Semi-parametric Bayesian Additive Regression Trees (spBART) with cross-validation for variable selection.
result Achieves strong out-of-sample discrimination (AUC = 0.96) in held-out validation set.
New method clusters high-dimensional data with anisotropic noise.
problem Clustering high-dimensional anisotropic mixtures with varying noise structures.
method Covariance Projected Spectral Clustering (COPO) method that projects data onto a low-dimensional space and reassigns clusters based on estimated covariances.
result COPO achieves minimax-optimal misclustering rates in Gaussian settings.
Paper solves a key problem in learning from high-dimensional covariance matrices.
problem Computing normalizing factors for Riemannian Gaussian distributions on high-dimensional covariance matrices.
method Equivalence with random matrix theory and log-normal matrix ensembles to approximate normalizing factors.
result Efficient approximation of normalizing factors with decreasing error as dimension increases.
New hierarchical model improves on standard practice for high-dimensional data.
problem Poor statistical performance in high-dimensional hierarchical models.
method Model effects as exchangeable across covariates and correlated across datasets.
result Empirical Bayes estimator outperforms classic approach in high-dimensional settings.
This paper presents a new method for estimating high dimensional covariance matrices. The method, permuted rank-penalized least-squares (PRLS), is based on a Kronecker product series expansion of the true covariance matrix. Assuming an i.i.d. Gaussian random sample, we establish high dimensional rates of convergence to…
Study high-dimensional covariance matrix estimators for complex portfolios, improving financial metrics.
problem Estimating covariance matrices in high-dimensional portfolios with nested and one-factor structures.
method Combining random matrix theory, free probability, deterministic equivalents, and two-step covariance estimators.
result Two-step estimators improve financial metrics in complex and one-factor covariance models.
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
New method speeds up sparse Bayesian learning without covariance matrix.
problem Sparse coding problem with uncertainty quantification.
method Covariance-free expectation maximization (CoFEM) that avoids explicit covariance matrix computation.
result Up to thousands of times faster than existing methods without sacrificing accuracy.
Model improves covariance estimation from shared and distinct datasets.
problem Limited sample sizes and shared covariance structure across related datasets.
method Spiked covariance model with shared subspace, closed-form pooling weight, and asymptotic guarantees.
result Improves estimation of high-dimensional covariance matrices from related datasets.
CR-FM-NES improves NES for high-dimensional optimization.
problem High-dimensional black-box optimization problems.
method CR-FM-NES extends FM-NES with a restricted covariance matrix representation.
result CR-FM-NES achieves significant speedup in high-dimensional problems.
New method for valid prediction sets in high-dimensional covariate shifts.
problem Valid prediction sets in high-dimensional covariate shifts.
method Likelihood-ratio regularized quantile regression (LR-QR) algorithm.
result LR-QR constructs valid prediction sets with desired coverage in target domain.
Develops inequalities for high-dimensional linear processes with dependent innovations.
problem Estimating high-dimensional VAR(p) systems and HAC covariance estimation.
method Concentration inequalities for l∞ norm of vector linear processes with sub-Weibull, mixingale innovations. result Obtained concentration bounds for the maximum entrywise norm of lag-h autocovariance matrices. Enhances power of covariance matrix tests for high-dimensional data.
problem Testing large covariance matrices in high-dimensional data.
method Proposes a new Fisher's combined probability test for quadratic form and maximum form statistics.
result Boosts power against more general alternatives.
Paper develops methods for estimating GLMs and SNR under proportional asymptotics.
problem Estimation of regression coefficients and SNR in high-dimensional GLMs.
method Method-of-Moments type estimators that bypass nuisance function estimation.
result Consistent and asymptotically normal estimators derived for targets of inference.
Overview of high-dimensional time series regression methods.
problem Estimation and inference with high-dimensional time series data.
method Limit theory for high-dimensional dependent data, asymptotic theory for time series regression, statistical learning methods.
result Main limit theory results and asymptotic theory for high-dimensional time series regression.
Spatially relaxed inference tackles high-dimensional linear models with correlated covariates.
problem Accurate inference is challenging in high-dimensional settings with spatially correlated covariates.
method Proposes ensembled clustered inference algorithms that control the δ-FWER under standard assumptions. result Ensembled clustered inference algorithms control the δ-FWER and achieve decent power. New method for cross-validation in high-dimensional data with dependent or heavy-tailed covariates.
problem Inconsistent cross-validation in high-dimensional settings with dependent or heavy-tailed covariates.
method ROTI-GCV framework for cross-validation under proportional asymptotics regime.
result Demonstrated accuracy of ROTI-GCV in synthetic and semi-synthetic settings.
Machine learning improves high-dimensional matrix estimation.
problem Efficient estimation of high-dimensional matrices.
method Integrates machine learning with classical optimization algorithms for high-dimensional matrix estimation.
result The reparameterized LADMM achieves faster convergence and higher accuracy.
Proposes a method to improve regression model performance with limited target data using fused-regularizer.
problem Model shifts and covariate shifts in high-dimensional regression.
method Two-step method with fused-regularizer to leverage source data for target task.
result Robust to covariate shifts, minimax-optimal under certain conditions, and validated by numerical tests.
The paper addresses the selection of synthetic data for improving classifier performance, focusing on the role of covariance shift.
problem The effectiveness of synthetic data in improving classifier performance is questioned, and the specific properties affecting this performance are unclear.
method The paper uses high-dimensional regression to analyze synthetic data selection, focusing on the covariance shift between synthetic and target distributions.
result The covariance shift between synthetic and target distributions affects the generalization error of classifiers, but the mean shift does not.
Adaptive classifier optimizes high-dimensional data with spiked covariance structure.
problem Classification of high-dimensional data with spiked covariance structure.
method Adaptive classifier that whitens data, screens features, and applies Fisher linear discriminant.
result The classifier is Bayes optimal under certain conditions and performs well on real and synthetic data.
Our article considers a Gaussian variational approximation of the posterior density in a high-dimensional state space model. The variational parameters to be optimized are the mean vector and the covariance matrix of the approximation. The number of parameters in the covariance matrix grows as the square of the number …
Better signal detection in undersampled data using joint and cross covariances.
problem Detecting shared signals in high-dimensional data with limited samples.
method Analysis of three covariance matrices: individual, cross, and joint.
result Joint and cross covariance matrices detect signals earlier than individual covariances.
New method optimizes model selection in high-dimensional regression models.
problem Model selection in high-dimensional misspecified regression models with covariate shift.
method Importance-weighted orthogonal greedy algorithm (IWOGA) and high-dimensional importance-weighted information criterion (HDIWIC).
result IWOGA + HDIWIC achieves optimal convergence rates in terms of prediction error.
Bayesian method uses data spectra to estimate non-sparse high-dimensional models.
problem Handling many parameters in high-dimensional Bayesian statistics.
method Data-adaptive Gaussian prior aligned with leading eigenvectors of sample covariance.
result Posterior contraction rates reveal the effect of spectral mass on prediction error.
New methods test correlation between network structure and node features.
problem Assessing correlation between network structure and node-level covariates.
method Four novel methods based on linear models and canonical correlation analysis.
result Theoretical guarantees and computational efficiency for testing network dependency.
Sliced inverse regression is a popular tool for sufficient dimension reduction, which replaces covariates with a minimal set of their linear combinations without loss of information on the conditional distribution of the response given the covariates. The estimated linear combinations include all covariates, making res…
We propose a simple imputation method for high-dimensional linear regression with missing data.
problem Handling missing covariates in high-dimensional linear regression.
method Impute missing entries with conditional mean of observed covariates and use standard LASSO or square-root LASSO.
result The imputation scheme retains minimax estimation rate and is pivotal for the square-root LASSO.
MediEncoder learns nonlinear representations for causal mediation analysis.
problem High-dimensional noisy covariates and mediators in biomedical studies.
method Coupled encoder-decoder architecture with cross-factor network.
result Improves estimation accuracy in high-dimensional causal mediation analysis.
New method stabilizes private LASSO for high-dimensional data with diverse covariate scales.
problem Privacy constraints and heterogeneity in covariate scales degrade LASSO stability and accuracy.
method Gram-based anisotropic objective perturbation to counteract covariate structure.
result Significantly improves convergence and statistical efficiency of private LASSO estimators.
Paper proposes a deep learning method for better covariance matrix forecasting.
problem Suboptimal predictive performance in traditional matrix volatility forecasting.
method Riemannian-geometry-aware deep learning framework for symmetric positive definite matrices.
result Our method outperforms traditional approaches in predictive accuracy.
Paper tackles imbalanced time series classification with a novel oversampling method.
problem Imbalanced time series classification challenges due to high dimensionality and correlation.
method Density-ratio based clustering followed by shrinkage technique for covariance estimation, then generating synthetic samples.
result OHIT outperforms state-of-the-art methods in F1, G-mean, and AUC metrics.
Proposes a deep learning framework for estimating counterfactual outcomes.
problem Challenges in estimating individual outcomes under different treatments.
method Deep variational Bayesian framework integrating factual and similar subjects' outcomes.
result Rigorously integrates individual features and similar subjects' responses for counterfactual outcomes.
We consider the estimation of integrated covariance (ICV) matrices of high dimensional diffusion processes based on high frequency observations. We start by studying the most commonly used estimator, the realized covariance (RCV) matrix. We show that in the high dimensional case when the dimension p and the observati…
Study examines robust regression in high dimensions with heavy-tailed data.
problem Analyzing robust regression in high-dimensional settings with heavy-tailed data.
method Sharp asymptotic characterisation of M-estimators and ridge regression in elliptical distributions.
result Ridge regression is optimal and universal for finite second moments but can decay faster without them.
Proposes a convex method to estimate GGMs with covariates.
problem Improving conditional independence structure estimation with covariates.
method Convex optimization framework for joint estimation of mean and precision matrix.
result Improved theoretical guarantees and practical utility demonstrated.
Meta-learning improves predictions with generalized ridge regression in high-dimensional settings.
problem Improving meta-learning performance in high-dimensional settings.
method Generalized ridge regression applied to high-dimensional multivariate random-effects linear models.
result Optimal predictive risk achieved when using the inverse of the covariance matrix of random coefficients.