Bayesian method finds patterns of mutual independence in data.
problem Investigating mutual independence in statistics.
method Bayesian model comparison using Markov chain Monte Carlo (MCMC).
result Automated search for patterns of mutual independence.
Extracts the finest pattern of mutual independence from data.
problem Inferring the finest mutual independence pattern from data.
method Estimate the set of valid patterns of dichotomic independence and use their intersection to infer the finest pattern.
result The method can estimate the finest mutual independence pattern from i.i.d. realizations of a multivariate normal distribution.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
New approach to disentangled representations using mutual information.
problem Disentangled representations lack sufficient inductive biases.
method Formulate disentanglement through mutual information and conditional independence.
result Violation of mutual information assumption leads to loss of disentanglement.
IndiSeek learns disentangled representations by balancing independence and completeness.
problem Learning disentangled representations with mutual information in multi-modal data.
method Combines independence-enforcing objective with a reconstruction loss that bounds conditional mutual information.
result Demonstrates effectiveness on synthetic data, CITE-seq, and real-world multi-modal benchmarks.
We apply both distance-based (Jin and Matteson, 2017) and kernel-based (Pfister et al., 2016) mutual dependence measures to independent component analysis (ICA), and generalize dCovICA (Matteson and Tsay, 2017) to MDMICA, minimizing empirical dependence measures as an objective function in both deflation and parallel m…
The paper studies how quickly samples from Langevin dynamics become independent.
problem Understanding the dependence between samples along Langevin dynamics and related algorithms.
method Measures dependence via Φ-mutual information and proves strong data processing inequalities. result The Φ-mutual information between samples decreases exponentially to zero. In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate Gaussian d…
This research designs a data-driven partition to test independence between continuous variables.
problem Testing independence between continuous random variables.
method Empirical log-likelihood statistic and data-driven tree-structured partition.
result Strongly consistent test of independence over probability families.
This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
We propose a test of independence of two multivariate random vectors, given a sample from the underlying population. Our approach, which we call MINT, is based on the estimation of mutual information, whose decomposition into joint and marginal entropies facilitates the use of recently-developed efficient entropy estim…
Neural network estimates mutual information for ICA.
problem Learning statistically independent features from noisy data.
method Minimizing mutual information estimated by a MINE network in a differentiable encoder network.
result Qualitatively equal solution to FastICA for blind-source-separation.
Paper benchmarks mutual info estimators on diverse distributions.
problem Evaluating mutual information estimators on complex, real-world distributions.
method Constructs a diverse family of known-ground truth distributions, proposes a benchmark platform.
result Highlights differences in classical and neural estimators' performance across various conditions.
The independence clustering problem is considered in the following formulation: given a set S of random variables, it is required to find the finest partitioning {U1,…,Uk} of S into clusters such that the clusters U1,…,Uk are mutually independent. Since mutual independence is the target, pairwise …
Self-Distilled Disentanglement improves counterfactual predictions by separating variables.
problem Improving counterfactual predictions in the presence of confounders and unobserved variables.
method Self-Distilled Disentanglement framework based on information theory.
result Effective counterfactual inference in synthetic and real-world datasets.
Cycles in causal learning cause feedback loops under intervention.
problem Cyclic causal structures lead to feedback loops in causal inference.
method Theoretical observations about self-referential distributions and their factorizations.
result Cyclic causal dependence can exist even when observational data suggest independence.
New test SCI for conditional independence on discrete data improves accuracy in causal discovery.
problem Testing conditional independence on discrete data often fails in practice.
method Proposes a new test based on stochastic complexity for discrete data.
result SCI is an asymptotically unbiased and L2 consistent estimator for conditional mutual information (CMI).
We propose a method for learning Markov network structures for continuous data without invoking any assumptions about the distribution of the variables. The method makes use of previous work on a non-parametric estimator for mutual information which is used to create a non-parametric test for multivariate conditional i…
OBF optimally filters features under independent Gaussian models.
problem Biomarker discovery from complex data.
method Optimal Bayesian feature selection under independent Gaussian models.
result OBF is consistent and optimal under mild conditions.
New framework improves multivariate time series forecasting by minimizing redundant information.
problem Improving multivariate time series forecasting with deep learning techniques.
method Cross-variable Decorrelation Aware feature Modeling (CDAM) and Temporal correlation Aware Modeling (TAM) to refine Channel-mixing and exploit temporal correlations.
result Significantly surpasses existing models in comprehensive tests.
Study reveals mutual information is crucial for understanding algorithm performance in stochastic convex optimization.
problem Uncertainty in capturing the exceptional performance of learning algorithms using existing information-theoretic generalization bounds.
method Examined the relationship between mutual information and generalization in stochastic convex optimization.
result Mutual information is necessary for true risk minimization in stochastic convex optimization, indicating existing bounds fall short.
Estimates conditional mutual information using a minmax formulation.
problem Estimating conditional mutual information in high dimensions.
method Uses a minmax optimization problem to train a neural network.
result Improves estimation accuracy compared to existing methods.
Meta-learning bound uses conditional mutual information.
problem Bounding generalization performance in meta-learning.
method Extends CMI framework to meta-learning with a meta-supersample.
result Explicit bound involving two CMI terms.
Half-AVAE enhances VAE for underdetermined ICA with adversarial training.
problem Challenges in ICA under underdetermined conditions.
method Encoder-free VAE with adversarial networks and EE terms.
result Half-AVAE outperforms baseline models in underdetermined ICA.
AutoTransfer improves subject transfer learning for biosignal datasets.
problem Subject transfer learning for challenging biosignal datasets.
method Regularization framework with mutual information or divergence penalties.
result Improves subject transfer learning performance on EEG, EMG, and ECoG datasets.
New estimator reduces bias and variance issues in mutual information estimation.
problem Difficulty in using variational MI estimators due to bias/variance tradeoffs and self-consistency issues.
method Developed a new estimator based on a unified perspective of variational approaches, focusing on variance reduction.
result Empirical results show improved bias-variance trade-offs compared to existing estimators.
This paper introduces efficient approximations for fairness criteria in regression models.
problem Measuring fairness in real-valued outcomes (regression settings) is computationally challenging.
method Fast approximations of mutual information for independence, separation, and sufficiency fairness criteria.
result The method achieves state-of-the-art accuracy/fairness tradeoffs in real-world datasets.
The paper explores tail diversification in financial markets using entropy and mutual information.
problem Tail diversification in financial time series.
method Statistical independence through differential entropy and mutual information, using moments as contrast functions.
result Tail covariance matrix is a key driver of tail diversification.
MSRL learns a representation maximizing mutual info with response variables.
problem Learning sufficient representations for complex, multi-dimensional data.
method Variational mutual information, deep neural networks, generalized Dudley's inequality.
result MSRL achieves consistent and accurate representation learning.
New method shows links can be braided open book bindings.
problem Understanding fibered links and their bindings.
method Mutual arc presentations and braided open books.
result Every fibered link is the binding of a braided open book.
Paper proposes a method to extract style features from unlabeled data.
problem Extracting fine-grained features like styles from unlabeled data.
method Contrastive conditioned variational autoencoders with mutual information constraints.
result The method efficiently extracts style features from real-world natural image datasets.
The paper introduces submodular information measures for machine learning applications.
problem Generalizing information-theoretic measures to non-random variables.
method Developing combinatorial information measures based on submodular functions.
result Submodular mutual information is submodular in one argument for certain submodular functions.
The conditional mutual information I(X;Y|Z) measures the average information that X and Y contain about each other given Z. This is an important primitive in many learning problems including conditional independence testing, graphical model inference, causal strength estimation and time-series problems. In several appl…
Paper proposes MIM-DRCFR to learn disentangled factors for better treatment effect estimation.
problem Learning disentangled factors precisely for individual-level treatment effect estimation.
method Multi-task learning framework with MI minimization criteria.
result MIM-DRCFR outperforms state-of-the-art methods in treatment effect estimation.
New framework IIA identifies innovations in general nonlinear vector autoregressive processes.
problem Limited generality of NVAR models due to additive innovation assumption.
method Independent Innovation Analysis (IIA) framework, assuming mutual independence and modulation by an auxiliary variable.
result Guarantees identifiability of innovations with arbitrary nonlinearities, up to permutation and component-wise invertible nonlinearities.
Feature selection methods are usually evaluated by wrapping specific classifiers and datasets in the evaluation process, resulting very often in unfair comparisons between methods. In this work, we develop a theoretical framework that allows obtaining the true feature ordering of two-dimensional sequential forward feat…
Mutual Information (MI) is often used for feature selection when developing classifier models. Estimating the MI for a subset of features is often intractable. We demonstrate, that under the assumptions of conditional independence, MI between a subset of features can be expressed as the Conditional Mutual Information (…
Estimators of information theoretic measures such as entropy and mutual information are a basic workhorse for many downstream applications in modern data science. State of the art approaches have been either geometric (nearest neighbor (NN) based) or kernel based (with a globally chosen bandwidth). In this paper, we co…
Boosting improves ICA for better component recovery.
problem Improving ICA's reliance on prior knowledge of sources.
method Maximizing likelihood via boosting and fixed-point unmixing.
result Boosting-based ICA outperforms existing methods.
We examine a class of deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight matrices are independent and ort…
In the field of machine learning, it is still a critical issue to identify and supervise the learned representation without manually intervening or intuition assistance to extract useful knowledge or serve for the downstream tasks. In this work, we focus on supervising the influential factors extracted by the variation…
MINDE estimates Mutual Information using neural diffusion models.
problem Estimating Mutual Information between random variables.
method Score-based diffusion models to estimate Kullback Leibler divergence.
result MINDE is more accurate than existing methods, especially for challenging distributions.
We present a variational renormalization group (RG) approach using a deep generative model based on normalizing flows. The model performs hierarchical change-of-variables transformations from the physical space to a latent space with reduced mutual information. Conversely, the neural net directly maps independent Gauss…
Multivariate pattern analyses approaches in neuroimaging are fundamentally concerned with investigating the quantity and type of information processed by various regions of the human brain; typically, estimates of classification accuracy are used to quantify information. While a extensive and powerful library of method…
Proposes a new method to interpret EEG classification models without needing a baseline.
problem Reliable interpretation of EEG classification models using integrated gradients.
method Compensated Integrated Gradients using Shapley sampling.
result The proposed method provides more reliable attributions than original integrated gradients.
Proposes Infomax and Domain-Independent Representations for robust causal inference.
problem Handling treatment selection bias and domain imbalance in causal inference with real-world data.
method Utilizes mutual information to learn domain-invariant representations that maximize predictive common information.
result Achieves state-of-the-art performance on causal effect inference across various data distributions.
Maximizes image representation dependence for self-supervised learning.
problem Learning meaningful image representations from unlabeled data.
method Maximizes Hilbert-Schmidt Independence Criterion (HSIC) between image transformations and identity.
result Matches state-of-the-art performance on ImageNet and other vision tasks.
Develops LSH schemes for f-divergences and mutual information loss.
problem Approximating nearest neighbors in high-dimensional probability distributions.
method General framework and specific LSH schemes for f-divergences and mutual information loss.
result Generalized Jensen-Shannon divergence can be approximated by Hellinger distance.