The paper identifies network bottlenecks using minimax paths in stochastic networks.
problem Identifying bottlenecks in networks with stochastic weights.
method Modeling as combinatorial semi-bandit problem, applying combinatorial Thompson Sampling, and approximating the original objective due to computational intractability.
result Established an upper bound on Bayesian regret and evaluated Thompson Sampling performance on real-world networks.
Develops a new unsupervised clustering method using Variational Information Bottleneck and Gaussian Mixture Model.
problem Unsupervised clustering of unlabeled data.
method Combines Variational Information Bottleneck and Gaussian Mixture Model in a deep neural network framework.
result Derives a new bound on the cost function and provides an algorithm for efficient computation.
Draft proposes adapting neural networks to match naive Bayes classifiers.
problem Bridge between neural networks and naive Bayes classifiers.
method Class-conditional compression and disentanglement using variational bounds.
result Latent representations enable naive Bayes classifier performance.
BCDE improves conditional density estimation for high-dimensional data.
problem High-dimensional conditional density estimation challenges.
method Hybrid training of CVAE with stochastic bottleneck.
result BCDE achieves competitive results in supervised and semi-supervised settings.
Proposes flexible auto-encoders for varying data dimensions.
problem Fixed latent dimensions limit data flexibility.
method Stochastic bottleneck with weighted dropouts.
result Seamless variable dimensionality reduction with high performance.
Proposes a new method to selectively access privileged information in reinforcement learning.
problem Selective compression of privileged information in reinforcement learning.
method Formulates a variational bandwidth bottleneck to decide stochastically whether to access privileged information.
result Improves generalization and reduces access to costly information in reinforcement learning experiments.
Improves deep network generalization for image sequence reconstruction.
problem Improving generalization of deep networks for inverse image reconstruction.
method Proposes a network optimized by a variational approximation of the information bottleneck principle with stochastic latent space.
result Demonstrates improved generalization ability of inverse reconstruction networks through stochasticity and information bottleneck.
Proposes a method to extract robust features that improve classifier robustness.
problem Improving classifier robustness to small perturbations in input space.
method Introduces an additional penalty term in the information bottleneck framework to minimize Fisher information, optimizing a variational bound using stochastic gradient descent.
result Optimally robust features are jointly Gaussian, and the method produces classifiers with increased robustness to perturbations.
New approach speeds up optimization with repeated gradient steps on same batch.
problem Performance bottlenecks in massive parallel pipelines with large batch sizes.
method Data echoing, taking repeated gradient steps on the same batch.
result Data echoing affords speedups on curvature-dominated part of convergence rate.
Survey on Information Bottleneck in Deep Learning.
problem Improving Deep Neural Networks (DNNs).
method Information Theory (IT) concepts applied to DNNs.
result IB is promising for analyzing and improving DNNs.
Introduces deterministic information bottleneck (DIB) replacing mutual information with entropy.
problem Compression and feature selection in lossy compression and clustering.
method Formulates deterministic information bottleneck (DIB) using entropy instead of mutual information, resulting in deterministic encoder (hard clustering).
result DIB outperforms IB in terms of DIB cost function and offers computational efficiency gains.
MFMs enable efficient reward alignment for generative models.
problem Computational bottleneck in controlling generative models.
method Meta Flow Maps (MFMs) extend consistency models and flow maps to stochastic regime for efficient value function estimation.
result MFMs enable inference-time steering and unbiased, off-policy fine-tuning to general rewards efficiently.
The paper proposes gradient sparsification to reduce communication costs in distributed optimization.
problem Reduction of communication overhead in distributed machine learning.
method Formulates a convex optimization problem to minimize gradient coding length, and proposes simple algorithms for approximate solution.
result The proposed sparsification techniques significantly reduce communication costs without sacrificing accuracy.
New research shows non-bottlenecked autoencoders can outperform bottlenecked ones for anomaly detection.
problem The necessity of a bottleneck in autoencoders for anomaly detection.
method Investigated two ways to remove bottlenecks: overparameterising the latent layer and introducing skip connections. Carried out extensive experiments on various AE types and datasets.
result Non-bottlenecked autoencoders can outperform bottlenecked ones, improving anomaly detection performance.
The paper interprets VQ-VAE loss as a form of information bottleneck.
problem Understanding the VQ-VAE loss function.
method Interpreted VQ-VAE loss as variational deterministic information bottleneck (VDIB) and variational information bottleneck (VIB).
result VQ-VAE loss can be derived from VDIB and approximated by VIB.
Proposes sigsoftmax to overcome the softmax bottleneck in language models.
problem Softmax function acts as a bottleneck in neural network representational capacity.
method Identifies the cause of softmax bottleneck and proposes sigsoftmax as a new activation function.
result Sigsoftmax outperforms softmax in language modeling tasks.
FrostNet improves INT8 quantization efficiency in mobile networks.
problem The importance of network architecture for optimal INT8 quantization.
method Quantization-aware training (QAT) with StatAssist and GradBoost, hardware-aware NAS.
result FrostNets achieve higher recognition accuracy with comparable latency when quantized.
Meta learning with information theory and Gaussian processes.
problem Few-shot learning problems.
method Information bottleneck, mutual information, variational approximations, Gaussian processes.
result Competitive accuracy on few-shot classification problems.
SCBMs model causal effects using low-dimensional bottlenecks.
problem Causal effect estimation in high-dimensional systems.
method Structural causal models with low-dimensional summary statistics.
result SCBMs provide a flexible framework for task-specific dimension reduction.
Wide neural networks with narrow bottlenecks behave like deep Gaussian processes.
problem Understanding the behavior of neural networks with narrow layers in the wide limit.
method Analyzing the wide limit of BNNs with narrow bottlenecks, showing they behave like a composition of GPs.
result Wide neural networks with narrow bottlenecks form a composition of GPs, termed a bottleneck NNGP.
BICompFL tackles bi-directional compression challenges in stochastic FL, reducing communication costs by an order of magnitude.
problem Communication bottleneck in federated learning, especially with stochastic updates.
method Introduces BICompFL, a bi-directional compression approach for stochastic federated learning.
result Significantly reduces communication costs (by an order of magnitude) while maintaining accuracy.
GLASS Flows improves flow and diffusion model performance by optimizing sampling efficiency.
problem Efficiency bottleneck in sampling Markov transitions for flow and diffusion models.
method Introduces GLASS Flows, a new sampling paradigm that simulates a 'flow matching model within a flow matching model' to sample Markov transitions efficiently.
result Eliminates the trade-off between stochastic evolution and efficiency in large-scale text-to-image models.
One-bit proximal method speeds up nonconvex stochastic optimization.
problem Reducing communication in distributed SGD for large datasets.
method Stochastic proximal gradient method using one-bit per update.
result The method achieves convergence rates similar to uncompressed SGD.
Study shows bottlenecks improve image segmentation quality.
problem Robust object discovery in real-world images remains challenging.
method Empirical investigation of reconstruction bottlenecks in GENESIS model.
result Reconstruction bottlenecks determine reconstruction and segmentation quality.
Paper revisits Deep Variational Information Bottleneck and proposes a new optimization approach.
problem Limitations of Deep Variational Information Bottleneck in optimizing mutual information.
method Proposes a new optimization approach by circumventing the limitation of requiring both Markov chains during optimisation.
result Shows how to optimise a lower bound for mutual information, circumventing the limitation of requiring both Markov chains.
Deep ResNets favor low bottleneck rank with proper hyperparameters.
problem Understanding the inductive bias of deep neural networks.
method Computed minimum-norm weights of a deep linear ResNet.
result Deep nonlinear ResNets have an inductive bias towards minimizing bottleneck rank.
Derives a new variational approach to information bottleneck.
problem Information theoretic inference and predictive modeling.
method Variational lower bound of predictive information bottleneck.
result Generalizes modern inference procedures and suggests new ones.
HSIC bottleneck trains deep networks without backpropagation.
problem Training deep neural networks with exploding and vanishing gradients.
method HSIC bottleneck, alternative to cross-entropy loss and backpropagation.
result HSIC bottleneck achieves comparable performance to backpropagation.
We propose a fast algorithm for spectral embedding using stochastic gradient descent.
problem Scalability issue in spectral embedding due to eigendecomposition bottleneck.
method Reformulate spectral embedding as a stochastic optimization problem, replacing orthogonality constraint with an orthogonalization matrix.
result Efficient algorithm based on mini-batch gradient descent that outperforms existing techniques in execution speed.
CB-APM uses analyst consensus as a bottleneck to interpret stock returns.
problem Tackles the challenge of understanding and predicting stock returns using professional beliefs.
method Embeds analyst consensus as a structural bottleneck, treating it as a sufficient statistic for market information.
result CB-APM portfolios exhibit strong monotonic return gradients and robust across different economic conditions.
Stochastic Gradient Descent (SGD) is a workhorse in machine learning, yet its slow convergence can be a computational bottleneck. Variance reduction techniques such as SAG, SVRG and SAGA have been proposed to overcome this weakness, achieving linear convergence. However, these methods are either based on computations o…
Wide neural networks become linear, but adding bottlenecks makes them bilinear or multilinear.
problem Understanding the transition of neural networks from linearity to higher-order functions.
method Analyzing the behavior of randomly initialized wide neural networks with and without bottleneck layers.
result Bottleneck layers transform the network's function from linear to bilinear or multilinear.
A deep RL method using a bottleneck state improves data efficiency.
problem Lack of data-efficiency in deep reinforcement learning.
method Model-based approach combining learned transition model and rollout simulations.
result The Bottleneck Simulator achieves excellent performance on natural language tasks.
New method quantifies redundant information using information bottleneck.
problem Quantifying redundant information among multiple sources.
method Formulated as an information bottleneck problem, termed redundancy bottleneck.
result Extracts information that best predicts the target without revealing source identity.
Improves GANs training through game theory.
problem Hard training of GANs due to antagonistic networks.
method Rewrote GAN training as a variational inequality and introduced a stochastic relaxed forward-backward algorithm.
result Algorithm converges to an exact solution or a neighborhood of it under monotonicity.
SySCD improves SCD scalability and speeds up training.
problem Scalability issues in parallel SCD algorithms.
method Developed a system-aware parallel SCD algorithm (SySCD) to avoid bottlenecks.
result Offers up to x42 speedup compared to state-of-the-art GLM solvers.
A new algorithm improves Bayesian federated learning by reducing communication overhead.
problem Bayesian federated learning constraints, including privacy, data ownership, and communication overhead.
method Proposes Quantised Langevin Stochastic Dynamics (QLSD) for Bayesian federated learning, using gradient compression and variance reduction techniques.
result Non-asymptotic and asymptotic convergence guarantees for QLSD and its improved versions.
The Mahler volume of a centrally symmetric convex body K is defined as M(K)= (Vol K)(Vol K^dual). Mahler conjectured that this volume is minimized when K is a cube. We introduce the bottleneck conjecture, which stipulates that a certain convex body K^diamond subset K X K^dual has least volume when K is an ellipsoid. If…
The Information bottleneck method is an unsupervised non-parametric data organization technique. Given a joint distribution P(A,B), this method constructs a new variable T that extracts partitions, or clusters, over the values of A that are informative about B. The information bottleneck has already been applied to doc…
Paper proposes a new coin betting method for training deep networks without learning rates.
problem Deep learning requires tuning many hyperparameters, especially learning rates.
method Reduces deep network training to a coin betting game, eliminating learning rates.
result Empirical and theoretical evidence shows the new method outperforms existing stochastic gradient algorithms.
Local AdaAlter reduces communication in SGD with adaptive learning rates.
problem Communication overhead in distributed training.
method Novel SGD variant with adaptive learning rates and reduced communication.
result Empirically reduces communication overhead by up to 30%.
MO2 learns useful behaviours from past experience for new tasks.
problem Discovering useful behaviours from past experience and transferring them to new tasks.
method Model-Based Offline Options (MO2) framework supporting sample-efficient bottleneck option discovery over continuous state-action spaces.
result MO2 outperforms recent option learning methods on complex long-horizon continuous control tasks.
Estimates log determinants using entropy for scalable machine learning.
problem Scalable calculation of matrix determinants is a bottleneck in machine learning.
method Maximum entropy framework with moment constraints for stochastic trace estimation.
result Significant improvement over state-of-the-art methods on various UFL sparse matrices.
VIB improves classification and uncertainty quantification.
problem Improving classification calibration and uncertainty quantification.
method Presented a simple case study of VIB.
result VIB improves classification calibration and uncertainty quantification naturally.
Unified approach for federated learning using MM optimization.
problem Scaling stochastic optimization to federated learning.
method Unified Majorize-Minimize (MM) framework for stochastic optimization, extended to federated learning.
result Unified algorithm \QSMM\ for federated learning that aggregates surrogate majorizing functions.
DPC uses physics and neural nets to solve SDEs.
problem Solving stochastic differential equations with missing physics.
method Physics-data fusion with conditional maximum mean discrepancy (CMMD) loss.
result DPC achieves highly accurate solutions on benchmark examples.
Graph neural networks struggle to propagate long-range information, causing over-squashing.
problem Graph neural networks struggle to propagate long-range information.
method Identified over-squashing as the bottleneck in GNNs, demonstrated on various models.
result Breaking the bottleneck improves GNNs' performance on long-range problems.
MPNNs struggle with class-bottlenecks and heterophily, leading to performance limitations.
problem Performance limitations of MPNNs under heterophily and structural bottlenecks.
method A statistical framework decomposing model performance into SNR components and proving bounds on sensitivity.
result Optimal graph structures for maximizing higher-order homophily are disjoint unions of single-class and two-class-bipartite clusters.