Tutorial on combining latent variable models with deep learning.
problem Combining latent variable models with deep learning to model natural language.
method Exploring variational inference to address intractable posterior inference and non-differentiability issues.
result Exploration of variational inference techniques to handle deep latent variable models.
New method identifies latent relationships in deep models without additional constraints.
problem Latent representations in deep latent variable models are not statistically identifiable.
method Identifies relationships between latent variables (distances, angles, volumes) under mild model conditions.
result Empirically demonstrates more reliable latent distances without additional labeled data.
Improved DSSMs for easier interpretable latent variables.
problem Complex and hard-to-interpret latent variables in DSSMs.
method Simplified predictive decoder and shrinkage priors.
result Interpretable latent variables improve forecasting performance.
Variational autoencoders learn deep latent models.
problem Learning deep latent-variable models.
method Principled framework using variational inference.
result Introduction to variational autoencoders and extensions.
Investigates latent variable models for useful generative concept representations.
problem Creating latent representations that support various concepts and attributes.
method Latent variable modeling, including latent variable models, latent representations, and latent spaces.
result Hierarchical latent representations and latent space vectors and geometry are effective for generative concept representations.
Langevin autoencoders improve deep latent variable models with efficient posterior sampling.
problem Efficient posterior sampling in deep latent variable models using MCMC.
method Amortized Langevin dynamics (ALD) replaces datapoint-wise sampling with encoder updates.
result ALD is valid as an MCMC algorithm with the target posterior as a stationary distribution.
Enhances deep kernel learning with stochastic latent variables for better model regularization.
problem Weak model regularization in deep kernel learning, especially on small datasets.
method Introduces DLVKL model with stochastic latent variables, NSDE for expressive posterior, and hybrid prior.
result DLVKL-NSDE outperforms existing deep GPs on large datasets.
Paper proposes a fast method for learning deep latent variable models.
problem Learning deep generative models with hierarchical latent variables.
method Noise initialized short run MCMC with variational optimization of step size.
result The method outperforms VAE in reconstruction and synthesis quality.
Proposes a non-parametric method for deep discrete latent variable models.
problem Learning sparse discrete latent representations in deep models.
method Iterative algorithm with Beta-Bernoulli process prior and local data scaling.
result Improves sparsity and scalability of deep discrete latent variable models.
BIVA uses a deep hierarchy of latent variables for better generative modeling.
problem Performance gap between VAE and autoregressive models in generative modeling.
method BIVA introduces a skip-connected generative model and a bidirectional stochastic inference path.
result BIVA reaches state-of-the-art test likelihoods and generates coherent images.
Method evaluates disentanglement in DLVMs, including those not aligned with latent axes.
problem Evaluate disentanglement in DLVMs, especially those not aligned with latent axes.
method Proposes a statistical method to discover generative factors of a dataset.
result Empirically demonstrates the advantage of the method on two datasets.
Unified framework for disentangling latent variables.
problem Unidentifiability of deep latent-variable models.
method Variational autoencoders and nonlinear ICA, with a factorized prior conditioned on an observed variable.
result Identification of true joint distribution over observed and latent variables is possible up to simple transformations.
The paper uses deep neural networks to estimate economic models without separability restrictions.
problem Estimating economic models with complex interaction effects and non-separable restrictions.
method Uses deep neural networks as a nonparametric sieve to approximate regression functions from nonlinear latent variable models.
result Economic shape, sparsity, or separability restrictions are imposed more straightforwardly when a flexible latent variable model is used.
A simple whitening method improves disentanglement in VAEs without sacrificing reconstruction quality.
problem Learning disentangled latent variables in VAEs.
method Proposes a whitening-based approach to disentangle latent variables in VAEs.
result The method finds interpretable latent factors without compromising reconstruction quality.
Proposes a deep latent variable model for MNAR data.
problem Missing data leading to biased results in MAR assumptions.
method Deep latent variable models with conditional no self-censoring.
result Establishes identifiability of MNAR data distribution.
This work studies the exact likelihood of DLVMs and its applications in inference.
problem The lack of attention to the exact likelihood of DLVMs and its implications for inference.
method Investigation of the properties of the exact likelihood, maximum likelihood estimation, and missing data imputation.
result The exact likelihood can be leveraged to ensure the existence of maximum likelihood estimates and improve missing data imputation.
New framework validates deep latent variable models for robustness.
problem Validating disentangled representations in neural networks.
method Causal perspective on representation learning, introducing a new metric for evaluation.
result New metric for evaluating deep latent variable models from labeled data.
Deep learning uses complex networks for high-dimensional data.
problem Computational inefficiency in training deep learning models.
method Use of hierarchical latent variables, efficient linear algebra, SGD optimization, and batch sampling.
result Efficient training and inference possible with optimized algorithms.
Enhanced VAE with DT improves flexibility in latent variable modeling.
problem Limitations of VAE's diagonal covariance matrix in matching true posterior distribution.
method Proposes dyadic transformation (DT) to model multivariate normal distributions.
result DT enhances posterior flexibility and achieves competitive results.
This work proposes a new method to train models with deep latent hierarchies using Optimal Transport.
problem Training models with deep latent hierarchies using VAEs often leads to the 'latent variable collapse' issue.
method Proposes a novel approach based on Optimal Transport to train models with deep latent hierarchies.
result The method avoids the 'latent variable collapse' issue and provides better sample generations and latent representation.
AEVB improves understanding of latent variable models.
problem Training latent variable models efficiently and understanding their limitations.
method Motivates AEVB from EM, emphasizing approximate E-step and M-step.
result AEVB tightens ELBO, improving model training.
Deep equilibrium models estimate latent variables from data.
problem Estimating latent variables from data.
method Generalized exponential family models, deep equilibrium networks.
result Deep equilibrium models solve MAP estimates for latent and transformation parameters.
Proposes SQUAD for better predictive uncertainty in deep latent models.
problem Intractable inference in deep latent variable models lead to overconfident predictions.
method Introduces Stochastic Quantized Activation Distributions (SQUAD) for flexible yet tractable latent variable distributions.
result The model provides competitive quality predictive uncertainty and learns non-linearities.
Advances deep latent variable models for more flexible text generation.
problem Limited representation power of VAEs due to Gaussian assumptions and posterior collapse.
method Develops sample-based variational distributions and an LVM to directly match aggregated posterior to prior.
result Demonstrates improved text generation in various scenarios.
Latent variable models improve RL by facilitating efficient learning and exploration.
problem Improving sample efficiency in reinforcement learning.
method Representation view of latent variable models for state-action value functions, incorporating kernel embeddings and UCB exploration.
result Established sample complexity of the proposed approach in online and offline settings, demonstrated superior performance in benchmarks.
Improved inference in probabilistic programs using attention mechanisms.
problem Inference failure in existing IC network architectures due to long-range dependency issues.
method Inference compilation with attention mechanism to model latent variables.
result Attention mechanism enhances proposal distributions to better match true posterior.
Bayesian methods enhance deep learning models by improving reliability and uncertainty.
problem Improving reliability and uncertainty awareness in deep learning models.
method Approximate Bayesian inference techniques, including SG-MCMC and VI, applied to deep learning models.
result Enhanced posterior inference for deep learning models, particularly in neural networks and generative models.
Deep models can't generate heavy-tailed samples well.
problem Understanding the limitations of deep generative models in generating samples with heavy tails.
method Unified framework using concentration of measure and convex geometry, Gromov-Levy inequality.
result Deep generative models are not universal generators and can only produce concentrated samples with light tails.
GEEN uses deep learning to estimate unobserved variables from observed data.
problem Estimating unobserved variables in latent variable models.
method GEEN uses deep learning with Kullback-Leibler distance to map observed measurements to latent variable realizations.
result GEEN provides a method to identify and estimate latent variables in a class of models.
Improved image reconstruction from sparse measurements using generative models.
problem Signal recovery from limited compressed measurements.
method Generative model with constrained latent variables for stable signal reconstruction.
result Improved reconstruction accuracy and preservation of realistic features.
Generative skip models avoid latent variable collapse in VAEs.
problem Latent variable collapse in VAEs prevents useful representations.
method Include skip connections in generative models to enforce strong links between latent variables and likelihood function.
result Generative skip models maintain similar predictive performance but reduce latent variable collapse and provide more meaningful representations.
Iterative models improve inference efficiency in deep latent variable models.
problem Inference models in deep latent variable models are computationally inefficient and have an amortization gap.
method Proposes iterative models that learn to perform inference optimization through repeated encoding of gradients.
result Iterative models outperform standard inference models on benchmark data sets of images and text.
Improved hierarchical discrete VAEs for better stability and performance.
problem Training stable and efficient hierarchical discrete VAEs with numerous latent variables.
method Introducing Relaxed-Responsibility Vector-Quantisation to parameterise discrete latent variables in a hierarchical structure.
result Achieved state-of-the-art bits-per-dim results for various standard datasets.
We describe \textit{deep exponential families} (DEFs), a class of latent variable models that are inspired by the hidden structures used in deep neural networks. DEFs capture a hierarchy of dependencies between latent variables, and are easily generalized to many settings through exponential families. We perform infere…
CRL uses causality to build interpretable AI models from complex data.
problem Interpreting deep neural networks' implicit representations.
method Causal representation learning (CRL) synthesizing latent variable models, causal graphical models, and nonparametric statistics.
result CRL can improve interpretability of generative AI models.
Algorithm learns latent variables for thermodynamically-consistent deep neural networks.
problem Predicting time evolution of large-scale physical systems with thermodynamic consistency.
method Sparse autoencoders and structure-preserving neural networks.
result Method conserves total energy and entropy inequality for both conservative and dissipative systems.
Improved generative models learn from unlabeled data with semi-supervised learning.
problem Improving generative accuracy from small labeled datasets.
method Developed a parameter-efficient deep semi-supervised generative model.
result Improved performance in disentangling latent variables and prediction.
Graphite learns graph node representations using deep latent variable models.
problem Learning graph node representations for machine learning tasks.
method Graphite uses deep latent variable generative models with graph neural networks and iterative graph refinement.
result Graphite outperforms other methods on tasks like density estimation, link prediction, and node classification.
This work extends identifiability analysis to sequential latent variable models, focusing on Switching Dynamical Systems.
problem Identifying latent variables in sequential data models.
method Proved identifiability of Markov Switching Models and established conditions for Switching Dynamical Systems.
result Identifiability of latent variables and non-linear mappings in Switching Dynamical Systems up to affine transformations.
Develops a method for lossless compression using latent variable models.
problem Lossless compression of large datasets.
method Bits back with asymmetric numeral systems (BB-ANS) using latent variable models.
result Achieves state-of-the-art lossless compression of full-size colour images.
Proposes a new model for unsupervised clustering with latent variables.
problem The challenge of unsupervised clustering in machine learning.
method Clustered Generator Model with continuous and discrete latent variables.
result Achieves competitive unsupervised clustering accuracy and disentangled latent representations.
New findings suggest latent regularization is unnecessary for high-quality image generation.
problem Improving image generation quality without latent regularization.
method Investigated the effect of latent regularization on image generation using learned priors.
result In the case of a sufficiently expressive prior, latent regularization is not necessary and may harm image quality.
A neural network finds causal relationships among latent variables.
problem Learning causal structure among latent variables in high-dimensional data.
method Redundant Input Neural Network (RINN) with modified architecture and regularized objective function.
result The RINN method successfully recovers latent causal structure between input and output variables.
Bayesian deep learning faces posterior collapse due to likelihood vs. prior competition.
problem Posterior collapse in Bayesian deep learning models.
method Identified competition between likelihood and prior regularization in a linear latent variable model.
result Posterior collapse is related to neural and dimensional collapse, suggesting a broader learning issue.
A recurring problem when building probabilistic latent variable models is regularization and model selection, for instance, the choice of the dimensionality of the latent space. In the context of belief networks with latent variables, this problem has been adressed with Automatic Relevance Determination (ARD) employing…
New method improves latent variable models without ignoring latent code.
problem Maximizing ELBO doesn't always yield good latent representations.
method Derive variational bounds on mutual information, derive rate-distortion curve, suggest new method.
result Demonstrates a family of models with identical ELBO but different characteristics.
Adapts VAE for survival analysis with missing data.
problem Inference on medical data sets with missing values and latent variables.
method Variational AutoEncoder (VAE) framework for survival analysis.
result Predicted distribution improves decision-making over classic models.
Modeling student course choices using latent variables.
problem Understanding student enrollment patterns in large universities.
method Probabilistic approach based on multilabel classification and mixture models.
result Demonstrated the model's ability to infer student interests guiding enrollment decisions.