Introduces a new geometric method for optimal experimental design.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
InfoOT improves data alignment by maximizing mutual information.
New bounds for optimal transport using Gaussian processes and rate-distortion functions.
Estimating mutual information is an important statistics and machine learning problem. To estimate the mutual information from data, a common practice is preparing a set of paired samples . However, in many situations, it…
Algorithm learns non-Gaussian graphical models via Hessian scores and triangular transport.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
Novel methods robustify Gromov-Wasserstein distance for cross-domain alignment.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
In this work, we develop a novel regularizer to improve the learning of long-range dependency of sequence data. Applied on language modelling, our regularizer expresses the inductive bias that sequence variables should have high mutual information even though the model might not see abundant observations for complex lo…
The identification of relevant features, i.e., the driving variables that determine a process or the properties of a system, is an essential part of the analysis of data sets with a large number of variables. A mathematical rigorous approach to quantifying the relevance of these features is mutual information. Mutual i…
We apply both distance-based (Jin and Matteson, 2017) and kernel-based (Pfister et al., 2016) mutual dependence measures to independent component analysis (ICA), and generalize dCovICA (Matteson and Tsay, 2017) to MDMICA, minimizing empirical dependence measures as an objective function in both deflation and parallel m…
The paper argues that normalized mutual information is biased in clustering and community detection.
Mutual information minimum spanning trees are used to explore nonlinear dependencies on Brazilian equity network in the periods from June/01/2015 to January/26/2016, in which Brazil was under the government of President Dilma Rousseff, and from January/27/2016 to September/08/2016 which includes the government transiti…
This paper presents a new methodology for clustering multivariate time series leveraging optimal transport between copulas. Copulas are used to encode both (i) intra-dependence of a multivariate time series, and (ii) inter-dependence between two time series. Then, optimal copula transport allows us to define two distan…
Neural estimator improves mutual information estimation in high dimensions.
Bounding the generalization error of learning algorithms has a long history, which yet falls short in explaining various generalization successes including those of deep learning. Two important difficulties are (i) exploiting the dependencies between the hypotheses, (ii) exploiting the dependence between the algorithm'…
Survey of Optimal Transport for model calibration.
New bounds for SGD generalize without mutual information terms.
Reshef et al. recently proposed a new statistical measure, the "maximal information coefficient" (MIC), for quantifying arbitrary dependencies between pairs of stochastic quantities. MIC is based on mutual information, a fundamental quantity in information theory that is widely understood to serve this need. MIC, howev…
InfoAtlas speeds up MI estimation for real-time data analysis.
Paper benchmarks mutual info estimators on diverse distributions.
Mutual information maximization has emerged as a powerful learning objective for unsupervised representation learning obtaining state-of-the-art performance in applications such as object recognition, speech recognition, and reinforcement learning. However, such approaches are fundamentally limited since a tight lower …
AMI framework improves text generation by optimizing mutual information between source and target.
SMOTE is one of the oversampling techniques for balancing the datasets and it is considered as a pre-processing step in learning algorithms. In this paper, four new enhanced SMOTE are proposed that include an improved version of KNN in which the attribute weights are defined by mutual information firstly and then they …
The paper studies how quickly samples from Langevin dynamics become independent.
This work improves independence tests for high-dimensional data.
In this paper we investigate the relationship between a general existence of transport maps of optimal couplings with absolutely continuous first marginal and the property of the background measure called essentially non-branching introduced by Rajala-Sturm (Calc.Var.PDE 2014). In particular, it is shown that the quali…
Recent works investigated the generalization properties in deep neural networks (DNNs) by studying the Information Bottleneck in DNNs. However, the mea- surement of the mutual information (MI) is often inaccurate due to the density estimation. To address this issue, we propose to measure the dependency instead of MI be…
We introduce a novel kernel that models input-dependent couplings across multiple latent processes. The pairwise joint kernel measures covariance along inputs and across different latent signals in a mutually-dependent fashion. A latent correlation Gaussian process (LCGP) model combines these non-stationary latent comp…
This study quantifies the scalability of k-Sliced Mutual Information (k-SMI) with dimension.
Improved bounds on learning algorithms' performance using conditional mutual information.
The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision of merging dependent multivariate normal variables in an AHC procedure as a Bay…
A novel Gromov-Wasserstein learning framework is proposed to jointly match (align) graphs and learn embedding vectors for the associated graph nodes. Using Gromov-Wasserstein discrepancy, we measure the dissimilarity between two graphs and find their correspondence, according to the learned optimal transport. The node …
The paper explores the relationship between joint mixability and negative dependence structures.
Improved method for encoding contingency tables reduces mutual information bias.
Enhances RJMCMC efficiency with non-linear transport-based proposals.
Sequence models assign probabilities to variable-length sequences such as natural language texts. The ability of sequence models to capture temporal dependence can be characterized by the temporal scaling of correlation and mutual information. In this paper, we study the mutual information of recurrent neural networks …
Estimates copula density for complex data distributions.
Maximizes image representation dependence for self-supervised learning.
New optimal transport method handles mass creation and destruction.
New bounds derived for machine learning algorithms using convex functions.
Framework uses optimal transport to quantify model risk in stochastic path laws.
Study of time-dependent metrics and connections in geometry.
Study reveals mutual information is crucial for understanding algorithm performance in stochastic convex optimization.
In this paper, we introduce and develop the theory of semimartingale optimal transport in a path dependent setting. Instead of the classical constraints on marginal distributions, we consider a general framework of path dependent constraints. Duality results are established, representing the solution in terms of path d…
Unified framework for robust, stable, and efficient density ratio estimation.
New bounds explain modern machine learning algorithms' generalization.
The paper proposes a method to detect and filter noisy or mislabeled data using pointwise mutual information.