Novel tests for genetic independence in high-dimensional data.
problem Testing independence in genetics studies with many variables.
method Defining premetric structures on genetic data support spaces.
result Solid theoretical framework and computationally-efficient implementations.
The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in …
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
Topological Pontryagin classes are algebraically independent in high-dimensional spaces.
problem Algebraic independence of topological Pontryagin classes.
method Analyzing rationalised cohomology of BTop(d).
result Topological Pontryagin classes are algebraically independent.
Develops a computationally tractable high-dimensional differential privacy estimator.
problem Differential privacy in high dimensions is computationally intractable.
method Combines high-dimensional robust statistics with differential privacy techniques.
result A computationally tractable algorithm with dimension-independent privacy loss.
Develops a nonparametric graphical model for conditional independence.
problem Evaluation of conditional independence without distributional assumptions.
method Nonlinear sufficient dimension reduction techniques applied to a nonparametric graphical model.
result Method outperforms existing methods in non-Gaussian settings and high-dimensional data.
We formulate and analyze a graphical model selection method for inferring the conditional independence graph of a high-dimensional nonstationary Gaussian random process (time series) from a finite-length observation. The observed process samples are assumed uncorrelated over time and having a time-varying marginal dist…
New method learns dependencies in high-dimensional data without graph assumptions.
problem Learning dependencies in nonparametric and high-dimensional settings.
method Neighbourhood lattice decomposition for nonparametric CI learning.
result Compact, non-graphical representation of CI exists in any graphical model.
Efficiently solves high-dimensional ODEs with probabilistic methods.
problem Solving high-dimensional ODEs with uncertainty quantification.
method Probabilistic numerical algorithm based on independence assumptions or Kronecker structure.
result Efficient probabilistic solutions for ODEs with millions of dimensions.
New method tests CMI using deep neural networks for high-dimensional data.
problem Testing conditional mean independence in high-dimensional settings.
method Population CMI measure and bootstrap-based testing with deep generative neural networks.
result Strong empirical performance and versatility in various scenarios.
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
Paper optimizes approximating high-dimensional diffusions by independent coordinates.
problem Optimizing approximations of high-dimensional diffusions by independent coordinates.
method Introduces independent projection as optimal for two criteria.
result Independent projection is optimal for two criteria related to entropy and convergence.
Improves joint distribution learning for high-dimensional datasets with complex correlations.
problem Conditional independence assumption limitations in VAE decoders for high-dimensional datasets.
method Cramer-Wold distance regularization and two-step learning method for flexible prior modeling.
result Effective joint distributional learning for high-dimensional datasets with multiple categorical variables.
Fitting high-dimensional data involves a delicate tradeoff between faithful representation and the use of sparse models. Too often, sparsity assumptions on the fitted model are too restrictive to provide a faithful representation of the observed data. In this paper, we present a novel framework incorporating sparsity i…
Self-Distilled Disentanglement improves counterfactual predictions by separating variables.
problem Improving counterfactual predictions in the presence of confounders and unobserved variables.
method Self-Distilled Disentanglement framework based on information theory.
result Effective counterfactual inference in synthetic and real-world datasets.
Study ridge regression for non-identically distributed data with varying variances.
problem Investigate high-dimensional regression with non-identical data variance.
method Propose a random effect model and use tools from random matrix theory.
result Highlight the double descent phenomenon in high-dimensional regression for certain variance profiles.
In data sets with many more features than observations, independent screening based on all univariate regression models leads to a computationally convenient variable selection method. Recent efforts have shown that in the case of generalized linear models, independent screening may suffice to capture all relevant feat…
Learning rate needs to decrease with higher data moments for effective ICA in high dimensions.
problem Slower convergence of ICA in high-dimensional data with high-order moments.
method High-dimensional ODE analysis of ICA algorithm under controlled moment structure.
result Critical learning rate threshold for effective ICA when moments are high.
Variable selection in high-dimensional space characterizes many contemporary problems in scientific discovery and decision making. Many frequently-used techniques are based on independence screening; examples include correlation ranking (Fan and Lv, 2008) or feature selection using a two-sample t-test in high-dimension…
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
New model separates images into independent factors quickly and easily.
problem Separating high-dimensional data like images into independent latent factors.
method Combines bijective feature maps with linear ICA model on the Stiefel manifold.
result Models converge quickly and achieve better unsupervised latent factor discovery.
High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.
problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.
We analyze the dynamics of an online algorithm for independent component analysis in the high-dimensional scaling limit. As the ambient dimension tends to infinity, and with proper time scaling, we show that the time-varying joint empirical measure of the target feature vector and the estimates provided by the algorith…
New method infers graph from dependent matrix data.
problem Inferring graph from dependent matrix data.
method Sparse-group lasso-based frequency-domain formulation with ADMM approach.
result Local convergence of inverse PSD estimators to true value.
Generalizes causal inference to high-dimensional outcomes.
problem Limited causal inference methods for multivariate outcomes.
method Formulates causal discrepancy tests for nominal variables, uses conditional independence tests.
result Causal CDcorr method improves finite sample validity and power.
Greedy selection works well in a toy model of independent increments.
problem Iterative selection of maximum-value processes from i.i.d. stochastic processes.
method Fixed greedy selection at each stage.
result Optimal strategy is greedy selection under independent increments.
Sparse graph learning for dependent time series using ADMM.
problem Inferring conditional independence graph of sparse, high-dimensional stationary multivariate Gaussian time series.
method Sparse-group lasso-based frequency-domain formulation and alternating direction method of multipliers (ADMM) optimization.
result Convergence of inverse PSD estimators to true value under certain conditions.
This paper presents the R package gRapHD for efficient selection of high-dimensional undirected graphical models. The package provides tools for selecting trees, forests and decomposable models minimizing information criteria such as AIC or BIC, and for displaying the independence graphs of the models. It has also some…
GIDS reduces high-dimensional response and predictor spaces, improving interpretability and computational efficiency.
problem Challenges in modeling interactions among high-dimensional multimodal data.
method Graph Independence Dual Screening (GIDS) framework that reduces both response and predictor dimensions.
result GIDS reduces feature space to 9,000 CpGs and 2,000 transcripts, revealing coordinated regulatory mechanisms.
Kernel methods and MLPs perform similarly to linear models in high dimensions.
problem Understanding the performance of kernel methods and MLPs in high-dimensional settings.
method Analysis of kernel methods and MLPs in a high-dimensional regime with proportional asymptotics.
result Linear models are optimal in high-dimensional settings when data is generated by kernel models with nonlinear relationships.
New bounds for MCMC on discrete spaces without dimension dependence.
problem High-dimensional statistical convergence analysis of MCMC methods.
method Combining multicommodity flow and single-element drift conditions.
result Informed Metropolis-Hastings algorithms achieve relaxation times independent of dimension.
rags2ridges simplifies graphical modeling of high-dimensional data.
problem Graphical modeling of high-dimensional precision matrices.
method Modular framework for extraction, visualization, and analysis of Gaussian graphical models.
result Provides a one-stop-shop for graphical modeling of high-dimensional precision matrices.
We study high-dimensional distribution learning in an agnostic setting where an adversary is allowed to arbitrarily corrupt an ε-fraction of the samples. Such questions have a rich history spanning statistics, machine learning and theoretical computer science. Even in the most basic settings, the only known…
Develops methods for estimating and providing confidence bands in sparse high-dimensional additive models.
problem Estimating and providing reliable confidence bands for nonparametric components in high-dimensional additive models.
method Integrates sieve estimation into a high-dimensional Z-estimation framework, employing a multiplier bootstrap procedure.
result Constructs uniformly valid confidence bands for the target component f1 in sparse high-dimensional additive models. The study uncovers the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
problem Understanding the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
method Extending the Convex Gaussian Min-Max Theorem to non-Gaussian settings, deriving asymptotic min-max characterizations, and proving asymptotic equivalence of regularizers.
result The projection of the ERM estimator onto a test covariate approximately follows a Gaussian convolution under certain conditions.
This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
Paper extends ICA to ISA with auxiliary variables for better speech representation learning.
problem Learning unsupervised speech representations with independent subspaces.
method Theoretical framework of nonlinear ISA with auxiliary variables.
result Proposes an algorithm to learn speech representations with independent subspaces.
We develop a mean-field theory for multi-component ICA in high dimensions.
problem Understanding multi-component ICA in high-dimensional settings.
method Asymptotically exact mean-field theory for multi-component online ICA.
result Explicit learnability boundaries and competition conditions linking step size, data moments, and initialization.
New method disentangles hidden data structures using HSIC and supervision.
problem Tackles the challenge of interpreting high-dimensional data.
method Supervised Independent Subspace Principal Component Analysis (sisPCA) using HSIC.
result Identifies and separates hidden data structures effectively.
In this paper, we present a novel framework incorporating a combination of sparse models in different domains. We posit the observed data as generated from a linear combination of a sparse Gaussian Markov model (with a sparse precision matrix) and a sparse Gaussian independence model (with a sparse covariance matrix). …
Estimates high-dimensional distributions using tree tensor networks.
problem Estimating high-dimensional probability distributions from i.i.d. samples.
method Tree-based tensor formats, empirical risk minimization, L2 contrast, orthogonal bases.
result Effective approximation of classical probabilistic models like Gaussian and graphical models.
LCIT tests conditional independence using latent representations.
problem Detecting conditional independencies in statistical and machine learning tasks.
method Generative framework for learning latent representations of target variables X and Y, then testing for remaining dependencies.
result LCIT outperforms state-of-the-art baselines consistently under different metrics and settings.
We propose a feature selection method that finds non-redundant features from a large and high-dimensional data in nonlinear way. Specifically, we propose a nonlinear extension of the non-negative least-angle regression (LARS) called N3LARS, where the similarity between input and output is measured through the norm…
New AI-block models for clustering high-dimensional variables based on maxima of random processes.
problem Clustering high-dimensional variables with weakly dependent maxima of random processes.
method Asymptotic Independent block (AI-block) models and an algorithm for variable clustering.
result The proposed AI-block models and algorithm can effectively identify clusters in high-dimensional data.
Gradient descent solves robust mean estimation in high dimensions.
problem High-dimensional robust mean estimation in the presence of adversarial outliers.
method Gradient descent with a structural lemma showing near-optimal solutions.
result Gradient descent can solve the robust mean estimation problem directly.
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
problem Testing conditional independence
method Testing-by-betting on an adaptively optimized Kernel Conditional Independence statistic
result Significantly reduces Type I error inflation while preserving high power
Functional magnetic resonance imaging (fMRI) produces data about activity inside the brain, from which spatial maps can be extracted by independent component analysis (ICA). In datasets, there are n spatial maps that contain p voxels. The number of voxels is very high compared to the number of analyzed spatial maps. Cl…