New measures quantify mutual dependence between multiple random vectors.
problem Measuring mutual dependence between multiple random vectors.
method Proposes three measures based on generalized distance covariance.
result Empirical and simplified empirical measures effectively test mutual independence.
Improves ICA via novel mutual dependence measures.
problem Improving Independent Component Analysis (ICA) for better component independence.
method Combines distance-based and kernel-based mutual dependence measures, introduces Latin hypercube sampling and Bayesian optimization for initialization.
result MDMICA outperforms other methods in terms of mutual independence of estimated components, especially when the ICA model is misspecified.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
A new regularizer boosts long-range dependency in sequence data.
problem Improving long-range dependency in sequence data models.
method Developed a mutual information regularizer to enhance sequence learning.
result The approach increases mutual information and likelihood on holdout data.
TCMI assesses mutual dependence of continuous variables without parametric assumptions.
problem Estimating mutual information from continuous distributions.
method TCMI extends mutual information to continuous variables using cumulative distributions.
result TCMI facilitates feature selection and ranking of variable sets.
Mutual info trees show higher risk in Brazilian equity network during transition.
problem Identifying nonlinear dependencies in Brazilian equity network.
method Used mutual information minimum spanning trees to compare with linear correlation.
result Mutual info trees indicate higher risk and power law tail in volatility transmission.
Combines chaining and mutual information methods for tighter generalization bounds.
problem Bounding generalization error of learning algorithms, especially in deep learning.
method Integrates chaining and mutual information methods to create a new generalization bound.
result Example shows significant improvement over existing bounds.
The paper argues that normalized mutual information is biased in clustering and community detection.
problem Bias in normalized mutual information for clustering and community detection.
method Introducing a modified version of mutual information to correct for information content and spurious dependence.
result The modified mutual information leads to different conclusions about which algorithms are best for community detection.
Wasserstein dependency measure improves unsupervised representation learning.
problem Incomplete representations from mutual information maximization.
method Wasserstein dependency measure using Wasserstein distance instead of KL divergence.
result Improved results on tasks with high mutual information.
This paper studies mutual information in sequence models and finds Transformers excel in capturing long-range dependencies.
problem Understanding the expressive power of sequence models in capturing temporal dependencies.
method Theoretical and empirical analysis of linear and nonlinear RNNs, including Transformers.
result Transformers can capture long-range mutual information more efficiently than RNNs.
Introduces a new geometric method for optimal experimental design.
problem Restrictive invariance properties of traditional OED approaches based on probability densities.
method Mutual transport dependence (MTD) using optimal transport theory.
result Demonstrates high-quality designs and flexibility compared to standard methods.
Neural estimator improves mutual information estimation in high dimensions.
problem Estimating mutual information in high dimensions is challenging.
method Parametrizing conditional densities with normalizing flows and using block autoregressive structure.
result Improved mutual information estimation on benchmark tasks.
A novel kernel models latent variable couplings across multiple processes.
problem Modeling latent variable couplings across multiple processes.
method Mutually-dependent Hadamard kernel and latent correlation Gaussian process (LCGP) model.
result The LCGP model recovers latent signal correlations and achieves state-of-the-art performance.
Empirical study on dependency in auto-encoding networks using HSIC.
problem Inaccurate measurement of mutual information in DNNs.
method Proposed using Hilbert-Schmidt Independence Criterion (HSIC) to measure dependency between layers in auto-encoding architectures.
result HSIC can measure dependence without density estimation, improving generalization evaluation.
Improved bounds for SGLD via data-dependent estimates.
problem Improving generalization bounds for noisy iterative learning algorithms.
method Variational characterization of mutual information and data-dependent priors.
result Significantly improved mutual information bounds for SGLD.
New bounds for SGD generalize without mutual information terms.
problem Generalizing SGD's learning dynamics for heavy-tailed distributions.
method Introducing a geometric decoupling term and bounding it computably.
result Proved generalization bounds without mutual information terms.
Reshef et al. recently proposed a new statistical measure, the "maximal information coefficient" (MIC), for quantifying arbitrary dependencies between pairs of stochastic quantities. MIC is based on mutual information, a fundamental quantity in information theory that is widely understood to serve this need. MIC, howev…
InfoAtlas speeds up MI estimation for real-time data analysis.
problem Efficiently measuring statistical dependency between high-dimensional datasets.
method Directly infers mutual information in a single forward pass using a pretrained model.
result Matches state-of-the-art accuracy with 100x speedup.
Paper benchmarks mutual info estimators on diverse distributions.
problem Evaluating mutual information estimators on complex, real-world distributions.
method Constructs a diverse family of known-ground truth distributions, proposes a benchmark platform.
result Highlights differences in classical and neural estimators' performance across various conditions.
AMI framework improves text generation by optimizing mutual information between source and target.
problem Previous MI approaches ignored the backward network, leading to loose variational bounds.
method AMI is a saddle point optimization framework that iteratively promotes and demotes generated instances.
result AMI significantly outperforms baselines on various text generation tasks.
The paper studies how quickly samples from Langevin dynamics become independent.
problem Understanding the dependence between samples along Langevin dynamics and related algorithms.
method Measures dependence via Φ-mutual information and proves strong data processing inequalities. result The Φ-mutual information between samples decreases exponentially to zero. This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
This study quantifies the scalability of k-Sliced Mutual Information (k-SMI) with dimension.
problem Understanding how SMI and its estimation rates depend on the ambient dimension.
method Developed k-SMI framework and derived bounds on MC estimates, established optimal convergence rates, and provided asymptotic results.
result Sharp bounds and optimal convergence rates for k-SMI estimation, revealing interplay with dimension and sample size.
Improved bounds on learning algorithms' performance using conditional mutual information.
problem Bounding the generalization error of learning algorithms.
method Introducing conditional mutual information and disintegrated mutual information to tighten bounds.
result New bounds are tighter than previous ones, especially for noisy, iterative algorithms.
The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision of merging dependent multivariate normal variables in an AHC procedure as a Bay…
Investigates mutual information for fitting deep models without hidden layer knowledge.
problem Deep nonlinear models' parameters are hard to fit due to hidden layer and non-affine relation.
method Used mutual information and KL divergence as objective functions for fitting models without hidden layer knowledge.
result Mutual information and KL divergence are successful objective functions for fitting deep models.
Improved method for encoding contingency tables reduces mutual information bias.
problem Mutual information bias in measuring label similarity.
method Improved method for encoding contingency tables to reduce information cost.
result Better bound on reduced mutual information in typical use cases.
Estimates copula density for complex data distributions.
problem Estimating copula density from observed data.
method Neural network-based copula density neural estimation (CODINE).
result Novel approach capable of modeling complex distributions.
Recent unsupervised representation learning methods maximize mutual information, but their success depends on architecture and estimator inductive biases.
problem Estimating mutual information is hard, and MI maximization can lead to entangled representations.
method The paper argues that the success of mutual information maximization methods depends on the choice of feature extractor architectures and the parametrization of MI estimators.
result The paper provides empirical evidence that the success of mutual information maximization methods is not solely due to MI properties.
Maximizes image representation dependence for self-supervised learning.
problem Learning meaningful image representations from unlabeled data.
method Maximizes Hilbert-Schmidt Independence Criterion (HSIC) between image transformations and identity.
result Matches state-of-the-art performance on ImageNet and other vision tasks.
New bounds derived for machine learning algorithms using convex functions.
problem Bounding generalization error in machine learning.
method Using strongly convex functions and subgaussian loss tails, derived new generalization bounds.
result Generalization bounds can be derived using any strongly convex function of the joint input-output distribution.
Study reveals mutual information is crucial for understanding algorithm performance in stochastic convex optimization.
problem Uncertainty in capturing the exceptional performance of learning algorithms using existing information-theoretic generalization bounds.
method Examined the relationship between mutual information and generalization in stochastic convex optimization.
result Mutual information is necessary for true risk minimization in stochastic convex optimization, indicating existing bounds fall short.
New bounds explain modern machine learning algorithms' generalization.
problem Explaining generalization behavior of modern machine learning algorithms.
method Proposes a new complexity measure based on empirical Rademacher complexity of an algorithm- and data-dependent hypothesis class.
result Obtains novel bounds with finite fractal dimension, simplifies proofs, and recovers known results.
The paper proposes a method to detect and filter noisy or mislabeled data using pointwise mutual information.
problem Detecting and filtering noisy or mislabeled data in deep learning models.
method A mutual information-based framework quantifying statistical dependencies between inputs and labels.
result The method effectively filters low-quality samples, improving classification accuracy by up to 15%.
New methods estimate point-wise dependency from neural MI models.
problem Estimating point-wise dependency between different events.
method Developed two methods: Probabilistic Classifier and Density-Ratio Fitting.
result Demonstrated effectiveness in MI estimation, self-supervised representation learning, and cross-modal retrieval.
MEG models for dynamic networks estimate dependencies and shared latent space relationships.
problem Modeling dynamic networks with shared latent space relationships and dependencies.
method MEG combines mutually exciting point processes and latent space models to estimate node-specific parameters and unobserved edges.
result MEG models can estimate intensities for unobserved edges, useful for anomaly detection in real-world applications.
A new MI estimator reduces complexity to linear time, achieving optimal MSE rates.
problem High computational complexity of MI estimators.
method Ensemble Dependency Graph Estimator (EDGE) combining LSH, dependency graphs, and ensemble bias-reduction.
result EDGE achieves optimal computational complexity O(N) and parametric MSE rate O(1/N). The paper proposes methods to extract and analyze individual variable information from complex dependencies.
problem Analyzing and understanding complex dependencies between multiple variables.
method Reversible normalization and iterative dependency reduction to extract individual information, and use it for direct mutual information and multi-feature Granger causality analysis.
result Decoupling of variables to analyze their individual information and direct mutual information transfers.
This paper uses SPk languages to explore LDD characteristics and multi-element dependencies.
problem Understanding and modeling Long Distance Dependencies (LDDs) with multi-element interactions.
method Generated datasets with various properties using Strictly k-Piecewise languages, analyzed using mutual information.
result The number of interacting elements in a dependency is a key characteristic of LDDs.
We discuss the connection between information and copula theories by showing that a copula can be employed to decompose the information content of a multivariate distribution into marginal and dependence components, with the latter quantified by the mutual information. We define the information excess as a measure of d…
A method to improve image synthesis diversity using mutual information.
problem Mode collapse in conditional GANs for multimodal image synthesis.
method Explicitly estimate and maximize mutual information between latent code and output image.
result Prevents mode collapse and encourages synthesis of diverse images.
The k-nearest neighbor classification method (k-NNC) is one of the simplest nonparametric classification methods. The mutual k-NN classification method (MkNNC) is a variant of k-NNC based on mutual neighborship. We propose another variant of k-NNC, the symmetric k-NN classification method (SkNNC) based …
We demonstrate that a popular class of nonparametric mutual information (MI) estimators based on k-nearest-neighbor graphs requires number of samples that scales exponentially with the true MI. Consequently, accurate estimation of MI between two strongly dependent variables is possible only for prohibitively large samp…
Generalizes underlap coefficient for multivariate group separation.
problem Quantifying distributional separation across groups in statistical learning.
method Generalizes underlap coefficient (UNL) to multivariate settings, studies its relationship with Bayes risk and mutual information, proposes an efficient importance sampling estimator.
result UNL as a measure of dependence between group labels and variables of interest, interpretable measure of partition-covariate dependence in clustering.
Is the large influence that mutual funds assert on the U.S. financial system spread across many funds, or is it is concentrated in only a few? We argue that the dominant economic factor that determines this is market efficiency, which dictates that fund performance is size independent and fund growth is essentially ran…
Method learns Markov networks from continuous data without distributional assumptions.
problem Learning Markov network structures for continuous data without distributional assumptions.
method Combines non-parametric mutual information estimator with constraint-based algorithm for learning graph structure.
result Shows superior structure learning accuracy compared to competing methods on synthetic data with non-linear dependencies.
Geometric estimator calculates dependency between multivariate samples.
problem Estimating dependency between multivariate samples.
method Randomly permuted minimal spanning tree over multivariate samples.
result Geometric Mutual Information (GMI) converges to a quantity equivalent to Henze-Penrose divergence.
A new test for conditional independence adapts to nonlinear dependencies efficiently.
problem Testing conditional independence in nonlinear and high-dimensional data.
method Nearest-neighbor estimator of conditional mutual information combined with local permutation scheme.
result The test reliably simulates null distribution and is better calibrated for non-smooth densities.