Paper revisits Deep Variational Information Bottleneck and proposes a new optimization approach.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we develop an unsupervised generative clustering framework that combines the Variational Information Bottleneck and the Gaussian Mixture Model. Specifically, in our approach, we use the Variational Information Bottleneck method and model the latent space as a mixture of Gaussians. We derive a bound on th…
In this paper, we provide an information-theoretic interpretation of the Vector Quantized-Variational Autoencoder (VQ-VAE). We show that the loss function of the original VQ-VAE can be derived from the variational deterministic information bottleneck (VDIB) principle. On the other hand, the VQ-VAE trained by the Expect…
New method improves deep neural network performance in regression tasks.
Interpretable machine learning has gained much attention recently. Briefness and comprehensiveness are necessary in order to provide a large amount of information concisely when explaining a black-box decision system. However, existing interpretable machine learning methods fail to consider briefness and comprehensiven…
In this draft, which reports on work in progress, we 1) adapt the information bottleneck functional by replacing the compression term by class-conditional compression, 2) relax this functional using a variational bound related to class-conditional disentanglement, 3) consider this functional as a training objective for…
CB-APM uses analyst consensus as a bottleneck to interpret stock returns.
CVIB uses information theory to learn counterfactuals from MNAR data without RCTs.
Deep learning networks have shown state-of-the-art performance in many image reconstruction problems. However, it is not well understood what properties of representation and learning may improve the generalization ability of the network. In this paper, we propose that the generalization ability of an encoder-decoder n…
In classic papers, Zellner demonstrated that Bayesian inference could be derived as the solution to an information theoretic functional. Below we derive a generalized form of this functional as a variational lower bound of a predictive information bottleneck objective. This generalized functional encompasses most moder…
The Information bottleneck method is an unsupervised non-parametric data organization technique. Given a joint distribution P(A,B), this method constructs a new variable T that extracts partitions, or clusters, over the values of A that are informative about B. The information bottleneck has already been applied to doc…
Proposes a multi-task learning model using variational information bottleneck.
We present a simple case study, demonstrating that Variational Information Bottleneck (VIB) can improve a network's classification calibration as well as its ability to detect out-of-distribution data. Without explicitly being designed to do so, VIB gives two natural metrics for handling and quantifying uncertainty.
New Variational InfoMax objective improves neural network performance.
This study analyzes VAEs using ID and II, revealing a transition in behaviour and distinct training phases.
Proposes a new method to selectively access privileged information in reinforcement learning.
Proposes a robust VIB approach using soft labels and mutual info estimation.
Unified information-theoretic objectives for training deep neural networks.
GeoIB uses information geometry to control compression in deep learning models.
Deep latent variable models are powerful tools for representation learning. In this paper, we adopt the deep information bottleneck model, identify its shortcomings and propose a model that circumvents them. To this end, we apply a copula transformation which, by restoring the invariance properties of the information b…
GIB improves neural network generalization by dynamically selecting task-relevant features across different sequential environments.
Tutorial on information bottleneck problems with connections to coding and learning.
In many applications, it is desirable to extract only the relevant aspects of data. A principled way to do this is the information bottleneck (IB) method, where one seeks a code that maximizes information about a 'relevance' variable, Y, while constraining the information encoded about the original data, X. Unfortunate…
VIB balances empirical and Bayesian approaches in predictive models.
Information Theory (IT) has been used in Machine Learning (ML) from early days of this field. In the last decade, advances in Deep Neural Networks (DNNs) have led to surprising improvements in many applications of ML. The result has been a paradigm shift in the community toward revisiting previous ideas and application…
We consider the problem of sufficient dimensionality reduction (SDR), where the high-dimensional observation is transformed to a low-dimensional sub-space in which the information of the observations regarding the label variable is preserved. We propose DVSDR, a deep variational approach for sufficient dimensionality r…
Meta learning with information theory and Gaussian processes.
New definition of disentanglement for non-independent factors of variation.
ResNet learns to compress information during training.
We introduce the HSIC (Hilbert-Schmidt independence criterion) bottleneck for training deep neural networks. The HSIC bottleneck is an alternative to the conventional cross-entropy loss and backpropagation that has a number of distinct advantages. It mitigates exploding and vanishing gradients, resulting in the ability…
Information bottleneck (IB) is a technique for extracting information in one random variable that is relevant for predicting another random variable . IB works by encoding in a compressed "bottleneck" random variable from which can be accurately decoded. However, finding the optimal bottleneck variab…
DAB enriches deep networks with uncertainty estimates using a codebook of training inputs.
New objective function improves model robustness.
This paper analyzes the training dynamics of binary neural networks using information bottleneck.
New IBOs resolve contradictions in mutual information optimization.
Zellner (1988) modeled statistical inference in terms of information processing and postulated the Information Conservation Principle (ICP) between the input and output of the information processing block, showing that this yielded Bayesian inference as the optimum information processing rule. Recently, Alemi (2019) re…
We propose a novel strategy for extracting features in supervised learning that can be used to construct a classifier which is more robust to small perturbations in the input space. Our method builds upon the idea of the information bottleneck by introducing an additional penalty term that encourages the Fisher informa…
Mutual information bounds generalization error in variational classifiers.
New model learns symmetry transformations from complex data.
Adversarial learning methods have been proposed for a wide range of applications, but the training of adversarial models can be notoriously unstable. Effectively balancing the performance of the generator and discriminator is critical, since a discriminator that achieves very high accuracy will produce relatively uninf…
New learning rules from information bottleneck improve deep learning without precise labels.
Domain adaptation aims to leverage the supervision signal of source domain to obtain an accurate model for target domain, where the labels are not available. To leverage and adapt the label information from source domain, most existing methods employ a feature extracting function and match the marginal distributions of…
Deep neural network learns discrete state abstractions for efficient planning.
New method uses cycle consistency to enforce invariance in latent space.
LogDet estimator improves entropy estimation in neural networks.
BF-VAE estimates uncertainty from LF and HF QoI samples.
HCBM improves deep learning explainability by non-linear concept aggregation.
Improves deep RL agents' generalization to unseen environments.