This work provides uncertainty intervals for semantic latent variables in disentangled latent spaces.
problem Challenges in providing meaningful uncertainty quantification for semantic information in disentangled latent spaces.
method Uses quantile regression to output heuristic uncertainty intervals, calibrates these intervals to contain true latent values, and propagates them through the generator.
result Reliably communicates semantically meaningful, principled, and instance-adaptive uncertainty in image super-resolution and image completion.
A game helps users understand latent factors in recommender systems.
problem Users struggle to understand the latent factors in recommender systems.
method Presented an output-agreement game to represent latent factors.
result Collected outputs reflect real-world characteristics of latent factors.
Identifies useful product reviews from online consumer feedback.
problem Finding useful reviews among noisy consumer feedback.
method Explores latent semantic factors in reviews using HMM-LDA model.
result Significant improvement in predicting useful reviews over baselines.
Plug-in method decomposes latent representations into interpretable factors.
problem Decomposing latent representations in neural networks without altering the original models.
method Factors' Decomposer-Entangler Network (FDEN) that learns to decompose latent representations into mutually independent factors.
result FDEN framework effectively decomposes latent representations into interpretable factors, maintaining original model integrity.
ML-VAE learns disentangled representations from grouped data.
problem Learning disentangled representations from grouped observations with minimal supervision.
method Multi-Level Variational Autoencoder (ML-VAE) that separates latent representation at group and observation levels.
result ML-VAE learns meaningful disentanglement of grouped data and enables manipulation of latent representation.
Unsupervised framework learns latent codes for controllable generation.
problem Challenging to achieve controllable generation with GANs.
method Self-training iterative feedback from discriminator to generator.
result Better disentanglement and semantic meaningful latent codes.
New method uses cycle consistency to enforce invariance in latent space.
problem Learning meaningful and independent factors of variation in datasets.
method Two separate latent subspaces, cycle consistency constraints, deep information bottleneck.
result Identifies more meaningful factors leading to sparser and interpretable models.
A new text clustering method using NMF and LSA improves stability and performance.
problem Text data's large, sparse term-document matrix makes clustering difficult.
method Proposes a new feature agglomeration method based on NMF and deterministic K-Means initialization.
result Significantly improves clustering performance and stability.
A method for user-controlled semantic image filling.
problem Generating coherent images with user-specified semantics.
method Deep generative model combining encoder, latent variables, and PixelCNN.
result User can control the inpainting of unobserved pixels while maintaining semantic coherence.
Linear Discriminant Analysis (LDA) is a well-known method for dimensionality reduction and classification. Previous studies have also extended the binary-class case into multi-classes. However, many applications, such as object detection and keyframe extraction cannot provide consistent instance-label pairs, while LDA …
Paper presents a framework for learning generative models with structured latent factors.
problem Learning controllable and generalizable representations of multivariate data with desired structural properties.
method The paper introduces a novel generative model framework that uses mask variables to model dependency structure and extends the multivariate information bottleneck theory.
result The framework learns semantically meaningful latent factors that reflect various desired structures and can automatically estimate dependency structure from data.
Paper proposes VAE-BPTF for better tensor factorization of sparse, imbalanced count data.
problem Inference of Bayesian Poisson-Gamma models for sparse and imbalanced count data is challenging.
method Variational auto-encoder framework with multi-layer perceptron networks for complex update information sharing and reweighting.
result VAE-BPTF outperforms current models in reconstruction errors and latent factor coherence across real-world datasets.
SAMI learns disentangled representations from data.
problem Learning disentangled representations from data.
method Combines diffusion models and VAEs to learn disentangled representations.
result SAMI learns disentangled representations that are interpretable and useful.
Improved genetic programming by optimizing mutation operators for continuous program search.
problem Small syntactic mutations in genetic programming can lead to unpredictable behavioral shifts.
method Learned a compact trading-strategy DSL, created a block-factorized embedding, and designed geometry-compiled mutation operators.
result Geometry-compiled mutation operators discover strong strategies using fewer evaluations and achieve higher Sharpe ratios.
Semantic TrueLearn uses semantic graphs to improve educational recommendation systems.
problem Challenges in handling semantic and hierarchical structure in knowledge areas.
method Introduces a novel learner model that exploits semantic relatedness between knowledge components using a Wikipedia link graph.
result Achieves statistically significant improvements in predictive performance for educational engagement.
Paper proposes structured semantic perturbations to improve adversarial attacks.
problem Vulnerability of deep neural networks to adversarial attacks.
method Manipulates semantic attributes via disentangled latent codes.
result Demonstrates the effectiveness of structured semantic perturbations.
TopicRNN integrates RNNs and latent topics for better semantic dependency capture.
problem Capturing long-range semantic dependencies in sequential data.
method End-to-end learned RNN with latent topics.
result TopicRNN outperforms existing contextual RNN baselines in word prediction and sentiment analysis.
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…
This work prevents variational autoencoders from collapsing by adding an auxiliary decoder.
problem Variational autoencoders can collapse into autodecoders, losing semantic information.
method Adding an auxiliary decoder to regularize the latent space.
result Auxiliary decoders increase semantic information in the latent space and reconstructions.
Representation learning systems typically rely on massive amounts of labeled data in order to be trained to high accuracy. Recently, high-dimensional parametric models like neural networks have succeeded in building rich representations using either compressive, reconstructive or supervised criteria. However, the seman…
New method neutralizes gender bias in word embeddings without losing semantic information.
problem Gender biases in word embeddings trained on human-generated corpora.
method Latent Disentanglement and Counterfactual Generation with siamese auto-encoder and gradient reversal layer.
result Our method outperforms existing debiasing methods in preserving semantic information and neutralizing gender biases.
In this tutorial, I will discuss the details about how Probabilistic Latent Semantic Analysis (PLSA) is formalized and how different learning algorithms are proposed to learn the model.
NMF with specific constraints is equivalent to LDA.
problem Dimensionality reduction of non-negative data.
method NMF with ℓ1 normalization constraints and Dirichlet prior. result NMF with these constraints is equivalent to LDA.
Enhances word embedding by transferring external knowledge.
problem Low-frequency words in semantic space.
method Latent Semantic Imputation (LSI) integrating graph theory and spectral embeddings.
result LSI generates reliable embedding vectors for low-frequency words.
Proposes CSG model to separate semantic and variation factors for OOD prediction.
problem Out-of-distribution examples cause conventional models to mix semantic and variation factors, leading to poor performance.
method Causal Semantic Generative model (CSG) based on causal reasoning, using variational Bayes for efficient learning and prediction.
result CSG can identify semantic factor and improve OOD prediction performance.
This paper presents the current state of a work in progress, whose objective is to better understand the effects of factors that significantly influence the performance of Latent Semantic Analysis (LSA). A difficult task, which consists in answering (French) biology Multiple Choice Questions, is used to test the semant…
CSTEM models document topics using VAE with semantic distance.
problem Inability of previous topic models to explain semantic relations correctly.
method Continuous semantic topic embedding model using variational autoencoder and Mahalanobis distance.
result Improves topic coherence and semantic relation explanation.
Method interprets GAN latent space via latent variable correlation analysis.
problem Understanding the inner workings of GANs.
method Analyzing correlation between latent variables and semantic contents in generated images.
result A method for controllable semantic content generation in GANs.
New framework aligns latent representations over-the-air using intelligent metasurfaces.
problem Heterogeneous transmitter-receiver models produce misaligned latent representations in semantic communication.
method Intelligent metasurfaces (SIM) emulate supervised and zero-shot semantic aligners directly in the wave domain.
result SIMs achieve up to 90% task accuracy in high SNR regimes, robust to low SNR.
The paper improves semantic interpolation in latent spaces of implicit models.
problem Interpolating between latent points in implicit models requires careful distributional matching.
method Proposes modifying the prior code distribution to concentrate more probability mass near the origin.
result Linear interpolation paths are shortest and pass through high-density regions, improving sample quality and semantics.
VJE learns latent representations without contrastive learning, providing probabilistic semantics.
problem Learning latent representations without contrastive signals.
method VJE maximizes a symmetric conditional evidence lower bound (ELBO) on paired encoder embeddings, using a Student-t distribution on a polar representation.
result VJE outperforms standard non-contrastive baselines in ImageNet-1K, CIFAR-10/100, and STL-10.
Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.
problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.
SEMASIA provides a large dataset of latent representations for model comparison.
problem Difficulty in comparing semantic structures across different neural network models.
method Collection of latent representations from 1700 pretrained models across various benchmarks.
result Consistent semantic organization across models and datasets.
New method controls generative models with continuous factors.
problem Lack of control over generative models and poor understanding of latent space.
method Introduces a method to find meaningful directions in latent space for precise control of generated images.
result Demonstrates effectiveness of method for GANs and variational auto-encoders.
CPFM integrates dimensionality reduction and reconstruction with flow networks.
problem Learning coupled continuous flows for data and embeddings.
method Coupled flow matching framework with Gromov-Wasserstein objective and dual-conditional flow network.
result CPFM preserves and recovers residual information in latent space.
Unified framework for disentangled VAEs improves latent space interpretability.
problem Challenges in evaluating and interpreting latent representations, especially for diverse data types.
method Unified bfVAE framework, FVH-LT, DBSR-LS, GAS, LSSI.
result bfVAE provides more favorable trade-off between disentanglement and reconstruction.
Mathematical models link perception and memory formation.
problem Linking perception and memory formation.
method Tensor decompositions and latent representations.
result Active semantic decoding process in perception.
PVAE learns disentangled representations from multimodal data.
problem Learning disentangled representations from multimodal sensory data.
method Partitioned Variational Autoencoder (PVAE) with multimodal generative model and training objectives.
result PVAE achieves over 99% accuracy on both modalities for semantic units.
The ability of the Generative Adversarial Networks (GANs) framework to learn generative models mapping from simple latent distributions to arbitrarily complex data distributions has been demonstrated empirically, with compelling results showing that the latent space of such generators captures semantic variation in the…
OrphicX generates causal explanations for GNNs by isolating latent causal factors.
problem Generating interpretable causal explanations for complex graph neural networks.
method Develops a generative model and objective function to isolate latent causal factors, maximizing information flow.
result OrphicX effectively identifies causal semantics, significantly outperforming alternatives.
New insights explain why β-VAEs fail at disentanglement.
problem Disentanglement performance of β-VAEs peaks at intermediate β and collapses as regularization increases. method Formalized information-theoretic mechanism, introduced λβ-VAE to stabilize disentanglement. result Strong regularization pressure leads to mutual information collapse in β-VAEs. A new method improves topic modeling accuracy using semantic filtering.
problem Improving topic modeling accuracy in text documents.
method Three-step process: generate word/word-pair, apply TF-IDF, merge similar semantic pairs.
result Improves topic accuracy by up to 12.99% compared to state-of-the-art models.
Chimera model combines link, content, and time for dynamic network analysis.
problem Community detection and prediction in evolving networks with dynamic changes.
method Shared factorization model that accounts for graph links, content, and temporal analysis.
result The approach simplifies temporal analysis and enables future community prediction.
New hierarchical VQ-VAE scheme improves image compression quality and features at low bitrates.
problem Low bitrate image compression maintaining quality and features.
method Hierarchical VQ-VAE with stochastic quantization and Markovian latent variables.
result High perceptual quality and semantic features at low bitrates.
Study finds whitepaper narratives do not predict market factor structure.
problem Predicting market behavior from cryptocurrency whitepaper claims.
method Zero-shot NLP classification combined with CP tensor decomposition of market data.
result Weak alignment between whitepaper claims and market statistics and latent factors.
Method infers domain-specific models without domain semantic descriptors.
problem Poor performance of standard supervised learning methods in unseen domains.
method Introduces latent domain vectors and neural networks for optimization.
result Inference of appropriate domain-specific models without semantic descriptors.
We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal mapping, or scaling, of power law distributed data. Power law distributed data are foun…
Paper reviews and compares NMF, PLSA, LBA, EMA, and LCA models.
problem Identifiability of latent models.
method Comparison and proof of identifiability.
result Identifiability of LBA, EMA, LCA, PLSA is unique if and only if NMF is unique.