New method controls posterior collapse in VAEs without network architecture constraints.
problem Posterior collapse in VAEs reduces diversity of generated samples.
method Introduces Latent Reconstruction (LR) loss to control posterior collapse.
result Controls posterior collapse on various datasets without architectural constraints.
Linear VAEs explain posterior collapse in VAEs via local maxima in log marginal likelihood.
problem Posterior collapse in VAEs where variational posterior matches prior for some latent variables.
method Analysis of linear VAEs and their relation to pPCA, proving ELBO does not introduce spurious local maxima.
result Linear VAEs have identifiable global maxima corresponding to principal component directions, explaining posterior collapse.
New research shows posterior collapse in VAEs isn't just about KL-divergence.
problem Posterior collapse in Variational Autoencoders (VAEs).
method Analyzes the loss surface of deep autoencoder networks and proves the existence of bad local minima.
result Posterior collapse in VAEs is caused by bad local minima, not just KL-divergence.
A simple pooling technique prevents posterior collapse in sequence VAEs.
problem Posterior collapse in sequence VAEs, causing models to ignore latent variables.
method Proposed a pooling technique to prevent posterior collapse.
result Significantly better data log-likelihood achieved compared to standard sequence VAEs.
New method prevents posterior collapse in iVAE models.
problem Posterior collapse in iVAE models where observations and ICs are independent given covariates.
method Developed CI-iVAE by considering a mixture of encoder and posterior distributions in the objective function.
result Prevents posterior collapse, resulting in latent representations with more information of the observations.
The variational autoencoder (VAE) is a popular combination of deep latent variable model and accompanying variational learning technique. By using a neural inference network to approximate the model's posterior on latent variables, VAEs efficiently parameterize a lower bound on marginal data likelihood that can be opti…
Bayesian deep learning faces posterior collapse due to likelihood vs. prior competition.
problem Posterior collapse in Bayesian deep learning models.
method Identified competition between likelihood and prior regularization in a linear latent variable model.
result Posterior collapse is related to neural and dimensional collapse, suggesting a broader learning issue.
Improved VAE models avoid posterior collapse in text modeling.
problem Posterior collapse in VAEs leads to poor data manifold parameterization.
method Coupled-VAE couples a VAE with a deterministic autoencoder to improve encoder and decoder parameterizations.
result Coupled-VAE consistently improves results in probability estimation and latent space richness.
This work tackles posterior collapse in conditional and hierarchical VAEs.
problem Posterior collapse in VAEs leads to poor latent variable representations.
method Theoretical analysis of linear conditional and hierarchical VAEs, empirical validation.
result Theoretical and empirical evidence of posterior collapse causes in conditional and hierarchical VAEs.
Variational autoencoders often collapse, showing latent variables are non-identifiable.
problem Posterior collapse in variational autoencoders due to non-identifiable latent variables.
method Proves latent variable non-identifiability causes posterior collapse. Proposes latent-identifiable models using Brenier maps and input convex neural networks.
result Latent-identifiable models resolve posterior collapse and provide meaningful representations.
New method controls posterior collapse in VAEs with theoretical guarantees.
problem Posterior collapse in VAEs where encoder ignores latent structure.
method Inverse Lipschitz constraint on decoder network.
result Controls degree of posterior collapse for various VAE models.
High-dimensional VAEs inevitably collapse to prior, requiring large datasets for good performance.
problem Posterior collapse in VAEs leads to poor representation learning quality.
method Analyzed a minimal VAE in a high-dimensional limit, evaluating conditions for posterior collapse with respect to beta and dataset size.
result VAEs face 'inevitable posterior collapse' beyond a certain beta threshold, regardless of dataset size.
New critics improve VAEs by preventing latent collapse.
problem Posterior collapse in VAEs where latent variables are ignored.
method Inference critics that detect and incentivize meaningful latent representations.
result Optimizing inference critics increases mutual information between latent and observations.
Levenshtein VAE prevents posterior collapse in text generation models.
problem Posterior collapse in VAEs where generators ignore latent variables.
method Replaces ELBO with a Levenshtein distance-based objective to prevent collapse.
result Levenshtein VAE produces more informative latent representations.
KL annealing helps VAEs avoid posterior collapse and overfitting.
problem Posterior collapse and overfitting in VAEs.
method Theoretical analysis of learning dynamics with KL annealing.
result Posterior collapse is inevitable when β exceeds a threshold. NE-VAE prevents posterior collapse in VAEs by embedding neighbors in latent space.
problem Posterior collapse in VAEs when strong decoders are used.
method Neighbor embedding in latent space to prevent collapse.
result NE-VAE produces qualitatively different latent representations with active latent dimensions.
When trained effectively, the Variational Autoencoder (VAE) is both a powerful language model and an effective representation learning framework. In practice, however, VAEs are trained with the evidence lower bound (ELBO) as a surrogate objective to the intractable marginal data likelihood. This approach to training yi…
Variational autoencoders learn distributions of high-dimensional data. They model data with a deep latent-variable model and then fit the model by maximizing a lower bound of the log marginal likelihood. VAEs can capture complex distributions, but they can also suffer from an issue known as "latent variable collapse," …
New objective function reduces posterior collapse in generative models.
problem Posterior collapse in generative models with small datasets and high latent dimensions.
method Replaces ELBO's KL divergence with MMD, introduces latent clipping.
result μ-VAE outperforms ELBO and β-VAE models in quality and stability.
Ricci flow smooths locally collapsing manifolds with controlled curvature.
problem Locally collapsing manifolds with controlled Ricci curvature.
method Ricci flow for a definite period of time, detecting collapsing infranil fiber bundles.
result Topological conditions detect collapsing infranil fiber bundles.
Molecule generation is to design new molecules with specific chemical properties and further to optimize the desired chemical properties. Following previous work, we encode molecules into continuous vectors in the latent space and then decode the vectors into molecules under the variational autoencoder (VAE) framework.…
HEBAE improves VAEs by adaptively balancing reconstruction and regularization.
problem Posterior collapse in VAEs leading to over-regularization and poor latent encoding.
method Hierarchical Empirical Bayes approach to probabilistic generative models.
result HEBAE generates higher quality samples with better FID scores.
Enhances Ricci flow theorem with scalar curvature bound.
problem Improving no-local-collapsing theorem of Ricci flow.
method Derives improved theorem under scalar curvature bound condition.
result Refines Perelman's no-local-collapsing theorem.
Stochastic variational inference for collapsed models has recently been successfully applied to large scale topic modelling. In this paper, we propose a stochastic collapsed variational inference algorithm for hidden Markov models, in a sequential data setting. Given a collapsed hidden Markov Model, we break its long M…
Deep latent variable models (LVM) such as variational auto-encoder (VAE) have recently played an important role in text generation. One key factor is the exploitation of smooth latent structures to guide the generation. However, the representation power of VAEs is limited due to two reasons: (1) the Gaussian assumption…
SentenceMIM learns rich latent representations for variable-length language data.
problem Challenges in learning VAEs for variable-length language data, especially posterior collapse.
method Probabilistic auto-encoder trained with Mutual Information Machine (MIM) learning.
result SentenceMIM learns informative latent representations with high mutual information.
LDReg addresses local dimensional collapse in self-supervised learning.
problem Local dimensional collapse in self-supervised learning representations.
method Local dimensionality regularization based on Fisher-Rao metric.
result LDReg improves representation quality and regularizes local and global dimensions.
A new method improves Bayesian filtering in nonlinear systems.
problem Bayesian filtering in nonlinear dynamical systems with non-Gaussian posteriors.
method Transport maps with block-triangular structure and gradient flows for MMD minimization.
result Accurate approximation of non-Gaussian posteriors without particle collapse.
We develop a framework for approximating collapsed Gibbs sampling in generative latent variable cluster models. Collapsed Gibbs is a popular MCMC method, which integrates out variables in the posterior to improve mixing. Unfortunately for many complex models, integrating out these variables is either analytically or co…
Bayesian approach reduces FL communication cost by one-shot.
problem High communication cost in optimization-based FL for high-dimensional models.
method Bayesian pseudocoresets and function-space inference for one-shot FL.
result Achieves prediction performance competitive to state-of-the-art with up to 2 orders of magnitude reduction in communication cost.
Due to the phenomenon of "posterior collapse," current latent variable generative models pose a challenging design choice that either weakens the capacity of the decoder or requires augmenting the objective so it does not only maximize the likelihood of the data. In this paper, we propose an alternative that utilizes t…
A new model tackles language generation issues by using discrete variational attention.
problem Information under-representation and posterior collapse in variational autoencoders.
method Proposes a discrete variational attention model with categorical distribution over attention mechanism.
result Enhances latent space for language generation and avoids posterior collapse.
Self-attention networks localize when eigenspectrum variance is small.
problem Self-attention mechanisms can lead to rank and entropy collapses, reducing model expressivity and trainability.
method Characterized attention localization using query-key eigenspectrum variance.
result Small eigenspectrum variance prevents both rank and entropy collapses, improving model performance.
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameter…
Topic models, and more specifically the class of Latent Dirichlet Allocation (LDA), are widely used for probabilistic modeling of text. MCMC sampling from the posterior distribution is typically performed using a collapsed Gibbs sampler. We propose a parallel sparse partially collapsed Gibbs sampler and compare its spe…
This paper reviews recent advancements in amortized Variational Inference.
problem Scalability and efficiency issues in traditional Variational Inference.
method Systematic review of various Variational Inference techniques, focusing on amortized approaches.
result Amortized Variational Inference improves scalability and efficiency for generative modeling tasks.
We study the geometry and topology of Riemannian 3-orbifolds which are locally volume collapsed with respect to a curvature scale. We show that a sufficiently collapsed closed 3-orbifold without bad 2-suborbifolds either admits a metric of nonnegative sectional curvature or satisfies Thurston's Geometrization Conjectur…
We localize the entropy functionals of G. Perelman and generalize his no-local-collapsing theorem and pseudo-locality theorem. Our generalization is technically inspired by further development of Li-Yau estimate along the Ricci flow. It can be used to show the Gromov-Hausdorff convergence of the Kähler Ricci flow on ea…
We prove that a 3-dimensional compact Riemannian manifold which is locally collapsed, with respect to a lower curvature bound, is a graph manifold. This theorem was stated by Perelman and was used in his proof of the geometrization conjecture.
Generalizes tools for studying collapsed manifolds to new geometry.
problem Studying collapsed manifolds with bounded sectional curvature.
method Generalizes fibration and stability theorems for compact group actions on manifolds with local bounded Ricci covering geometry.
result Two generalized results used in Xiaochun Rong's work on almost flat manifolds.
Topology of non-orientable spaces without boundary is studied.
problem Topology of non-collapsed RCD spaces without boundary.
method Studied the stability of non-orientability and topology under Gromov-Hausdorff convergence.
result Non-orientable spaces without boundary have a stable ramified double cover.
Study shows local topologies of certain geometric spaces.
problem Local topological properties of geometric spaces.
method Analysis of Gromov-Hausdorff limits of manifolds with bounded Ricci curvature.
result Local b1 vanishes for regular loci in limits of non-collapsed manifolds. Variational language models seek to estimate the posterior of latent variables with an approximated variational posterior. The model often assumes the variational posterior to be factorized even when the true posterior is not. The learned variational posterior under this assumption does not capture the dependency relat…
We study collapsed manifolds with Ricci bounded covering geometry i.e., Ricci curvature is bounded below and the Riemannian universal cover is non-collapsed or consists of uniform Reifenberg points. Via Ricci flows' techniques, we partially extend the nilpotent structural results of Cheeger-Fukaya-Gromov, on collapsed …
A new method improves Bayesian deep learning by balancing scalability and accuracy.
problem Scalability issues in Bayesian neural networks.
method Collapsed inference scheme that performs Bayesian model averaging using collapsed samples.
result Significant improvements over existing methods in predictive performance and uncertainty estimation.
This paper analyzes VAE approximation errors in conditional exponential families.
problem Posterior collapse and approximation errors in VAEs.
method Analysis of ELBO objective and conditional exponential families.
result The ELBO optimizer pulls away from the likelihood optimizer towards a consistent subset of models.
For regular particle filter algorithm or Sequential Monte Carlo (SMC) methods, the initial weights are traditionally dependent on the proposed distribution, the posterior distribution at the current timestamp in the sampled sequence, and the target is the posterior distribution of the previous timestamp. This is techni…
This paper computes exact posterior distributions of mixture weights in hierarchical Bayesian models.
problem Uncertainty in class membership or data-generating processes in heterogeneous data.
method Exact marginalization of mixture weights using dynamic programming and FFT for two components, and joint dynamic program for K >= 3 components.
result Exact posterior distributions of mixture weights are finite mixtures of Beta distributions, providing credible intervals and per-observation local false-discovery rates.