New method controls posterior collapse in VAEs without network architecture constraints.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Posterior collapse in Variational Autoencoders (VAEs) arises when the variational posterior distribution closely matches the prior for a subset of latent variables. This paper presents a simple and intuitive explanation for posterior collapse through the analysis of linear VAEs and their direct correspondence with Prob…
Variational autoencoders (VAEs) hold great potential for modelling text, as they could in theory separate high-level semantic and syntactic properties from local regularities of natural language. Practically, however, VAEs with autoregressive decoders often suffer from posterior collapse, a phenomenon where the model l…
New method prevents posterior collapse in iVAE models.
The variational autoencoder (VAE) is a popular combination of deep latent variable model and accompanying variational learning technique. By using a neural inference network to approximate the model's posterior on latent variables, VAEs efficiently parameterize a lower bound on marginal data likelihood that can be opti…
Bayesian deep learning faces posterior collapse due to likelihood vs. prior competition.
Improved VAE models avoid posterior collapse in text modeling.
In narrow asymptotic settings Gaussian VAE models of continuous data have been shown to possess global optima aligned with ground-truth distributions. Even so, it is well known that poor solutions whereby the latent posterior collapses to an uninformative prior are sometimes obtained in practice. However, contrary to c…
This work tackles posterior collapse in conditional and hierarchical VAEs.
Variational autoencoders often collapse, showing latent variables are non-identifiable.
New method controls posterior collapse in VAEs with theoretical guarantees.
High-dimensional VAEs inevitably collapse to prior, requiring large datasets for good performance.
New critics improve VAEs by preventing latent collapse.
Levenshtein VAE prevents posterior collapse in text generation models.
KL annealing helps VAEs avoid posterior collapse and overfitting.
NE-VAE prevents posterior collapse in VAEs by embedding neighbors in latent space.
Variational autoencoders learn distributions of high-dimensional data. They model data with a deep latent-variable model and then fit the model by maximizing a lower bound of the log marginal likelihood. VAEs can capture complex distributions, but they can also suffer from an issue known as "latent variable collapse," …
When trained effectively, the Variational Autoencoder (VAE) is both a powerful language model and an effective representation learning framework. In practice, however, VAEs are trained with the evidence lower bound (ELBO) as a surrogate objective to the intractable marginal data likelihood. This approach to training yi…
Ricci flow smooths locally collapsing manifolds with controlled curvature.
Molecule generation is to design new molecules with specific chemical properties and further to optimize the desired chemical properties. Following previous work, we encode molecules into continuous vectors in the latent space and then decode the vectors into molecules under the variational autoencoder (VAE) framework.…
HEBAE improves VAEs by adaptively balancing reconstruction and regularization.
Enhances Ricci flow theorem with scalar curvature bound.
Stochastic variational inference for collapsed models has recently been successfully applied to large scale topic modelling. In this paper, we propose a stochastic collapsed variational inference algorithm for hidden Markov models, in a sequential data setting. Given a collapsed hidden Markov Model, we break its long M…
Deep latent variable models (LVM) such as variational auto-encoder (VAE) have recently played an important role in text generation. One key factor is the exploitation of smooth latent structures to guide the generation. However, the representation power of VAEs is limited due to two reasons: (1) the Gaussian assumption…
LDReg addresses local dimensional collapse in self-supervised learning.
A new method improves Bayesian filtering in nonlinear systems.
We develop a framework for approximating collapsed Gibbs sampling in generative latent variable cluster models. Collapsed Gibbs is a popular MCMC method, which integrates out variables in the posterior to improve mixing. Unfortunately for many complex models, integrating out these variables is either analytically or co…
Bayesian approach reduces FL communication cost by one-shot.
Due to the phenomenon of "posterior collapse," current latent variable generative models pose a challenging design choice that either weakens the capacity of the decoder or requires augmenting the objective so it does not only maximize the likelihood of the data. In this paper, we propose an alternative that utilizes t…
A new model tackles language generation issues by using discrete variational attention.
Self-attention networks localize when eigenspectrum variance is small.
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameter…
Topic models, and more specifically the class of Latent Dirichlet Allocation (LDA), are widely used for probabilistic modeling of text. MCMC sampling from the posterior distribution is typically performed using a collapsed Gibbs sampler. We propose a parallel sparse partially collapsed Gibbs sampler and compare its spe…
This paper reviews recent advancements in amortized Variational Inference.
We study the geometry and topology of Riemannian 3-orbifolds which are locally volume collapsed with respect to a curvature scale. We show that a sufficiently collapsed closed 3-orbifold without bad 2-suborbifolds either admits a metric of nonnegative sectional curvature or satisfies Thurston's Geometrization Conjectur…
We prove that a 3-dimensional compact Riemannian manifold which is locally collapsed, with respect to a lower curvature bound, is a graph manifold. This theorem was stated by Perelman and was used in his proof of the geometrization conjecture.
We localize the entropy functionals of G. Perelman and generalize his no-local-collapsing theorem and pseudo-locality theorem. Our generalization is technically inspired by further development of Li-Yau estimate along the Ricci flow. It can be used to show the Gromov-Hausdorff convergence of the Kähler Ricci flow on ea…
Topology of non-orientable spaces without boundary is studied.
Study shows local topologies of certain geometric spaces.
Variational language models seek to estimate the posterior of latent variables with an approximated variational posterior. The model often assumes the variational posterior to be factorized even when the true posterior is not. The learned variational posterior under this assumption does not capture the dependency relat…
We study collapsed manifolds with Ricci bounded covering geometry i.e., Ricci curvature is bounded below and the Riemannian universal cover is non-collapsed or consists of uniform Reifenberg points. Via Ricci flows' techniques, we partially extend the nilpotent structural results of Cheeger-Fukaya-Gromov, on collapsed …
A new method improves Bayesian deep learning by balancing scalability and accuracy.
This paper analyzes VAE approximation errors in conditional exponential families.
For regular particle filter algorithm or Sequential Monte Carlo (SMC) methods, the initial weights are traditionally dependent on the proposed distribution, the posterior distribution at the current timestamp in the sampled sequence, and the target is the posterior distribution of the previous timestamp. This is techni…
This paper computes exact posterior distributions of mixture weights in hierarchical Bayesian models.
We will simplify the earlier proofs of Perelman's collapsing theorem of 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's semi-convex analysis of distance functions to construct the desired local Seifert fibration structure on collapsed 3-manifolds. The verification of Perelma…
SentenceMIM is a probabilistic auto-encoder for language data, trained with Mutual Information Machine (MIM) learning to provide a fixed length representation of variable length language observations (i.e., similar to VAE). Previous attempts to learn VAEs for language data faced challenges due to posterior collapse. MI…
Improved co-clustering for robust data analysis.