New metrics for high-dimensional data improve on energy distance.
problem Testing equality of distributions and independence in high dimensions.
method Proposed new metrics inheriting properties of energy distance and others.
result Improved metrics detect homogeneity and independence in high dimensions.
Modeling data as being sampled from a union of independent subspaces has been widely applied to a number of real world applications. However, dimensionality reduction approaches that theoretically preserve this independence assumption have not been well studied. Our key contribution is to show that 2K projection vect…
Random Fourier Features reduce kernel matrix reconstruction error without dimensionality dependence.
problem Error reduction in kernel matrix reconstruction for high-dimensional data.
method Random Fourier Features with theoretical error bounds.
result Error probability is independent of data dimensionality.
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
Develops a nonparametric graphical model for conditional independence.
problem Evaluation of conditional independence without distributional assumptions.
method Nonlinear sufficient dimension reduction techniques applied to a nonparametric graphical model.
result Method outperforms existing methods in non-Gaussian settings and high-dimensional data.
Novel tests for genetic independence in high-dimensional data.
problem Testing independence in genetics studies with many variables.
method Defining premetric structures on genetic data support spaces.
result Solid theoretical framework and computationally-efficient implementations.
Topological Pontryagin classes are algebraically independent in high-dimensional spaces.
problem Algebraic independence of topological Pontryagin classes.
method Analyzing rationalised cohomology of BTop(d).
result Topological Pontryagin classes are algebraically independent.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.
Ultrahigh-dimensional variable selection plays an increasingly important role in contemporary scientific discoveries and statistical research. Among others, Fan and Lv [J. R. Stat. Soc. Ser. B Stat. Methodol. 70 (2008) 849-911] propose an independent screening framework by ranking the marginal correlations. They showed…
New method learns dependencies in high-dimensional data without graph assumptions.
problem Learning dependencies in nonparametric and high-dimensional settings.
method Neighbourhood lattice decomposition for nonparametric CI learning.
result Compact, non-graphical representation of CI exists in any graphical model.
The number of functionally independent scalar invariants of arbitrary order of a generic pseudo--Riemannian metric on an n--dimensional manifold is determined.
Paper optimizes approximating high-dimensional diffusions by independent coordinates.
problem Optimizing approximations of high-dimensional diffusions by independent coordinates.
method Introduces independent projection as optimal for two criteria.
result Independent projection is optimal for two criteria related to entropy and convergence.
The n-dimensional torus is uniquely characterized by specific harmonic forms.
problem Characterizing the n-dimensional torus via harmonic forms.
method Analyzing closed 1-forms on the torus to determine unique properties.
result The n-dimensional torus is the unique manifold supporting a linearly independent set of (n-1) closed 1-forms whose product determines a non-zero cohomological class.
FIT is a fast nonparametric test for conditional independence.
problem Testing conditional independence for high-dimensional data.
method Based on the conditional independence principle, FIT assesses whether additional variables improve predictions.
result FIT is significantly faster and more accurate than existing methods for large datasets.
New knots have unique Whitehead doubles.
problem Understanding independence in knot concordance groups.
method Using 4-dimensional constructions and Whitehead doubles.
result Infinite families of knots with independent Whitehead doubles.
A three-dimensional closed orientable orbifold (with no bad suborbifolds) is known to have a geometric decomposition from work of Perelman along with earlier work of Boileau-Leeb-Porti and Cooper-Hodgson-Kerckhoff. We give a new, logically independent, unified proof of the geometrization of orbifolds, using Ricci flow.…
New method for analyzing complex data spaces.
problem Dimensionality reduction and learning data representations for continuous spaces.
method Manifold factorization based on spectral graph methods.
result Recovering factors yields meaningful lower-dimensional representations.
A frame independent formulation of analytical mechanics in the Newtonian space-time is presented. The differential geometry of affine values i.e., the differential geometry in which affine bundles replace vector bundles and sections of one dimensional affine bundles replace functions on manifolds, is used. Lagrangian a…
Conditional independence testing is an important problem, especially in Bayesian network learning and causal discovery. Due to the curse of dimensionality, testing for conditional independence of continuous variables is particularly challenging. We propose a Kernel-based Conditional Independence test (KCI-test), by con…
Develops a computationally tractable high-dimensional differential privacy estimator.
problem Differential privacy in high dimensions is computationally intractable.
method Combines high-dimensional robust statistics with differential privacy techniques.
result A computationally tractable algorithm with dimension-independent privacy loss.
The k-dimensional coding schemes refer to a collection of methods that attempt to represent data using a set of representative k-dimensional vectors, and include non-negative matrix factorization, dictionary learning, sparse coding, k-means clustering and vector quantization as special cases. Previous generalizat…
We formulate and analyze a graphical model selection method for inferring the conditional independence graph of a high-dimensional nonstationary Gaussian random process (time series) from a finite-length observation. The observed process samples are assumed uncorrelated over time and having a time-varying marginal dist…
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
New method tests CMI using deep neural networks for high-dimensional data.
problem Testing conditional mean independence in high-dimensional settings.
method Population CMI measure and bootstrap-based testing with deep generative neural networks.
result Strong empirical performance and versatility in various scenarios.
IMA addresses non-identifiability in nonlinear ICA by assuming orthogonal Jacobian columns.
problem Non-identifiability in nonlinear ICA.
method IMA assumes orthogonal Jacobian columns and extends to manifold settings.
result IMA circumvents non-identifiability issues and can be beneficial for higher-dimensional observations.
Greedy selection works well in a toy model of independent increments.
problem Iterative selection of maximum-value processes from i.i.d. stochastic processes.
method Fixed greedy selection at each stage.
result Optimal strategy is greedy selection under independent increments.
In this paper we investigate the relationship between the existence of parallel semi-Riemannian metrics of a connection and the reducibility of the associated holonomy group. The question as to whether the holonomy group necessarily reduces in the presence of a specified number of independent parallel semi-Riemannian m…
Deep nets learn structured densities without dimensionality issues.
problem Learning structured densities in high dimensions.
method Simple L2-minimizing loss for neural networks. result Dimension-independent convergence rates for neural networks.
CIRCE measures conditional independence for learning invariant features.
problem Learning invariant features while being conditionally independent of a distractor.
method CIRCE is a measure of conditional independence applied as a regularizer in feature learning.
result CIRCE provides a zero value if and only if features are conditionally independent of the distractor given the target.
An entirely new and independent enumeration of the crystallographic space groups is given, based on obtaining the groups as fibrations over the plane crystallographic groups, when this is possible. For the 35 ``irreducible'' groups for which it is not, an independent method is used that has the advantage of elucidating…
This work explores the connection between distances and kernels for conditional independence.
problem Measuring conditional independence in various fields like causal discovery and feature selection.
method Investigates the relationship between conditional independence measures induced by distances and reproducing kernels.
result Some kernel-based conditional independence measures are not equivalent to distance-based measures.
In data sets with many more features than observations, independent screening based on all univariate regression models leads to a computationally convenient variable selection method. Recent efforts have shown that in the case of generalized linear models, independent screening may suffice to capture all relevant feat…
Variable selection in high-dimensional space characterizes many contemporary problems in scientific discovery and decision making. Many frequently-used techniques are based on independence screening; examples include correlation ranking (Fan and Lv, 2008) or feature selection using a two-sample t-test in high-dimension…
Fitting high-dimensional data involves a delicate tradeoff between faithful representation and the use of sparse models. Too often, sparsity assumptions on the fitted model are too restrictive to provide a faithful representation of the observed data. In this paper, we present a novel framework incorporating sparsity i…
Efficiently solves high-dimensional ODEs with probabilistic methods.
problem Solving high-dimensional ODEs with uncertainty quantification.
method Probabilistic numerical algorithm based on independence assumptions or Kronecker structure.
result Efficient probabilistic solutions for ODEs with millions of dimensions.
Bayesian structure learning for high-dimensional data using recursive bootstrap.
problem Bayesian structure learning for domains with hundreds of variables.
method Non-parametric bootstrap, recursive structure learning, combining bootstrap with constraint-based learning.
result The proposed method learns better MAP models and more reliable causal relationships than other state-of-the-art methods.
A new test for conditional independence adapts to nonlinear dependencies efficiently.
problem Testing conditional independence in nonlinear and high-dimensional data.
method Nearest-neighbor estimator of conditional mutual information combined with local permutation scheme.
result The test reliably simulates null distribution and is better calibrated for non-smooth densities.
High-dimensional ICA analysis shows asymptotic decoupling and PDE solutions.
problem Understanding the dynamics of high-dimensional ICA in the scaling limit.
method Analysis of an online ICA algorithm in the high-dimensional scaling limit, showing convergence to a PDE.
result The time-varying joint empirical measure converges to a PDE solution representing the algorithm's performance.
LCIT tests conditional independence using latent representations.
problem Detecting conditional independencies in statistical and machine learning tasks.
method Generative framework for learning latent representations of target variables X and Y, then testing for remaining dependencies.
result LCIT outperforms state-of-the-art baselines consistently under different metrics and settings.
Generative diffusion models gradually memorize training data, losing independent dimensions.
problem Understanding how generative diffusion models memorize training data, especially on low-dimensional manifolds.
method Measuring latent dimensionality via the learned score field, proposing a geometric memorization theory.
result Generative diffusion models experience a smooth collapse of their capacity to vary across independent directions as data become scarce, leading to near point-wise replication of salient features.
Paper discovers simplicial complexes connecting trained models for improved ensembling.
problem Improving robustness and accuracy of deep learning ensembles.
method Identifies mode-connecting simplicial complexes on loss surfaces.
result Efficiently builds simplicial complexes for ensembling, outperforming independent ensembles.
Study torsion's impact on 2D affine Killing vectors on homogeneous surfaces.
problem Effects of torsion on affine Killing vectors on homogeneous surfaces.
method Complete description of Lie algebras of affine Killing vector fields on homogeneous surfaces.
result Complete description of Lie algebras of affine Killing vector fields on homogeneous surfaces.
Improves joint distribution learning for high-dimensional datasets with complex correlations.
problem Conditional independence assumption limitations in VAE decoders for high-dimensional datasets.
method Cramer-Wold distance regularization and two-step learning method for flexible prior modeling.
result Effective joint distributional learning for high-dimensional datasets with multiple categorical variables.
New model separates images into independent factors quickly and easily.
problem Separating high-dimensional data like images into independent latent factors.
method Combines bijective feature maps with linear ICA model on the Stiefel manifold.
result Models converge quickly and achieve better unsupervised latent factor discovery.
BBVI converges nearly dimensionally independent for log-concave targets.
problem Efficiently optimizing variational parameters in high-dimensional spaces.
method Proved convergence rate of BBVI with reparametrization gradient for log-concave targets.
result BBVI converges with nearly independent dimension dependence for log-concave targets.
This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
Functional magnetic resonance imaging (fMRI) produces data about activity inside the brain, from which spatial maps can be extracted by independent component analysis (ICA). In datasets, there are n spatial maps that contain p voxels. The number of voxels is very high compared to the number of analyzed spatial maps. Cl…