We propose a parsimonious topic model for text corpora. In related models such as Latent Dirichlet Allocation (LDA), all words are modeled topic-specifically, even though many words occur with similar frequencies across different topics. Our modeling determines salient words for each topic, which have topic-specific pr…
ETM identifies field-specific keywords in text classification.
problem Unsupervised text classification with field-specific keywords.
method Weighted Lasso penalty and pairwise Kullback-Leibler divergence penalty for topic separation.
result ETM improves topic coherence by 22% and 10% compared to LDA.
Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…
New method learns more diverse topics from text documents considering paragraph structure.
problem Classic Topic Models ignore word position and use symmetric priors, limiting topic diversity.
method Exploits paragraph structure to distinguish between general and specific topics.
result Shows improved topic diversity and relevance in structured documents.
A new method for topic detection using hierarchical latent tree models.
problem Hierarchical topic detection in document collections.
method Graphical models (HLTMs) with binary variables at different levels representing word co-occurrence patterns and document clusters.
result Captures both general and specific topics at different levels of a hierarchical structure.
System identifies related passages for easier text comprehension.
problem Expensive and limited cross-reference resources for complex texts.
method Fine-grained topic modeling to find topically related verse pairs.
result System can produce cost-effective cross-references.
Self-supervised learning excels in topic modeling by being less model-specific.
problem How self-supervised learning discovers useful representations in topic models.
method Applying self-supervised learning objectives to topic model-generated data.
result Self-supervised learning objectives can recover useful posterior information for topic models, outperforming misspecified models.
New method predicts ICU stay for pancreatitis patients.
problem Predicting ICU stay for pancreatitis patients.
method Survival-supervised topic modeling with elastic-net regularized Cox model and anchor words.
result Our method is as accurate as best baselines but more interpretable.
Paper proposes a method to monitor research topic evolution.
problem Difficulty in tracking research topic diffusion and evolution.
method Deep Non-negative Autoencoder with information divergence measurement.
result Identifies evolution of research topics and discovers topic diffusions.
Improved neural topic model for semi-supervised learning.
problem Representing textual data in an interpretable manner with limited labeled data.
method Label-Indexed Neural Topic Model (LI-NTM) that combines deep generative models with semi-supervised learning.
result LI-NTM outperforms existing models in document reconstruction and classifier performance.
Hybrid approach combines topic and graph embeddings for legal document clustering.
problem Challenges in classifying legal texts due to domain-specific language and limited labeled data.
method Combines unsupervised topic and graph embeddings with a supervised model.
result Improves clustering quality over text-only or graph-only embeddings.
This paper extracts topic-specific subgraphs from Wikidata for easier research.
problem Time-consuming and task-specific tuning in using Wikidata for graph processing.
method Unified framework to extract topic-specific subgraphs from Wikidata.
result Standardized sub-graphs for easier evaluation of graph processing algorithms.
A new neural topic model using optimal transport improves document representation and topic coherence.
problem Challenges in achieving good document representation and coherent/diverse topics in existing NTMs.
method Proposes a neural topic model via optimal transport, learning topic distribution by minimising OT distance to document word distributions.
result Significantly outperforms state-of-the-art NTMs on discovering coherent and diverse topics.
Proposes a new deep topic model using MBN and Lasso.
problem Difficult optimization problem in topic modeling.
method Multilayer bootstrap network (MBN) for dimension reduction, supervised Lasso for topic word discovery.
result Effectiveness demonstrated on 20-newsgroups and TDT2 corpora.
This paper improves topic modeling by embedding words and topics together.
problem Topic models struggle with short documents and approximate inference.
method Model each document as a mixture of word embeddings and each topic as a mixture of topic embeddings.
result The method optimizes topic embeddings to minimize semantic differences between words and topics.
Study finds common poetic themes across languages over time.
problem Understanding thematic evolution in different poetic traditions.
method Applied Latent Dirichlet Allocation (LDA) to poetry corpora of four languages.
result Identified common themes and their temporal trends across poetic traditions.
Recent advances in topic models have explored complicated structured distributions to represent topic correlation. For example, the pachinko allocation model (PAM) captures arbitrary, nested, and possibly sparse correlations between topics using a directed acyclic graph (DAG). While PAM provides more flexibility and gr…
CorEx learns topics without assumptions, incorporating human input.
problem Complexity and detailed assumptions in topic modeling.
method Information-theoretic framework, anchor words for minimal human input.
result Topics comparable to LDA with minimal human intervention.
Analyzes Twitter users' opinions on self-driving cars.
problem Understanding public perception of self-driving cars.
method Annotated Twitter dataset, topic modeling, sentiment classification using Twitter features.
result People are generally optimistic but also concerned about self-driving cars.
Bayesian dynamic topic model improves topic prevalence prediction.
problem Estimating document-specific topic proportions in dynamic topic models.
method Developed a Bayesian dynamic topic model with covariates and dynamic structure, including polynomial trends and periodicity. Used MCMC algorithm with Polya-Gamma data augmentation and Gaussian approximation.
result Explicitly modeling polynomial and periodic behavior improves topic prevalence prediction.
The paper investigates topic models, ensuring their statistical identifiability and accuracy.
problem Lack of formal theoretical investigation of topic model identifiability and estimation accuracy.
method Proposes a maximum likelihood estimator (MLE) based on integrated likelihood, introducing new geometric identifiability conditions.
result Introduces weaker conditions for topic model identifiability, allowing a broader investigation.
Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can help us cluster a huge collection into different topics or find a subset of the co…
New Gamma-Poisson model improves topic selection for short text.
problem Topic modelling for short text using Poisson distribution.
method Gamma-Poisson mixture model with collapsed Gibbs sampler.
result Gamma-Poisson model selects more accurate number of topics.
We develop a nested hierarchical Dirichlet process (nHDP) for hierarchical topic modeling. The nHDP is a generalization of the nested Chinese restaurant process (nCRP) that allows each word to follow its own path to a topic node according to a document-specific distribution on a shared tree. This alleviates the rigid, …
Improved BERT model with latent persona and topic variables.
problem Improving BERT's domain-specific utility while maintaining generalization.
method Combining BERT with Universal Transformer, adding latent persona and topic variables.
result Pre-trained model for social texts outperforms baseline.
GraphSTONE uses topic models to capture graph structures, improving GCN performance.
problem GCNs focus too much on node features and not enough on graph structures.
method GraphSTONE employs topic models of graphs to capture structural topics, which guide the aggregation of node features.
result GraphSTONE outperforms GCNs in performance, efficiency, and interpretability.
Traditional Relational Topic Models provide a way to discover the hidden topics from a document network. Many theoretical and practical tasks, such as dimensional reduction, document clustering, link prediction, benefit from this revealed knowledge. However, existing relational topic models are based on an assumption t…
Incorporating the side information of text corpus, i.e., authors, time stamps, and emotional tags, into the traditional text mining models has gained significant interests in the area of information retrieval, statistical natural language processing, and machine learning. One branch of these works is the so-called Auth…
Study reveals AI's spontaneous topic changes in text prediction.
problem AI's inability to spontaneously switch topics like humans.
method Defined topic as Token Priority Graphs (TPGs) and analyzed self-attention models.
result AI can only switch topics if lower-priority tokens outnumber higher-priority ones.
Efficient nonparametric Bayesian topic model for social media text.
problem Text analytics for social media data.
method Hierarchical Pitman-Yor processes for topic modeling.
result Nonparametric model outperforms existing parametric models.
Neural model predicts survival outcomes and reveals feature relationships.
problem Predicting time-to-event outcomes and understanding feature relationships in clinical data.
method Survival and topic modeling combined in a neural network framework.
result Neural survival-supervised topic models achieve competitive accuracy with interpretability.
Paper presents a multi-label topic model for financial texts with high performance and insights into market reactions.
problem Analyzing financial text data for market reactions and understanding topic interactions.
method Trained a multi-label topic model on a financial text database, achieved high macro F1 score, and investigated topic interactions.
result Model achieves high performance (macro F1 > 85%) and reveals significant market reactions to topic co-occurrences.
Paper proposes a new model to infer topic-based connection structures from noisy adjacency matrices.
problem Traditional community detection assumes static connection structures, but real-world connections vary based on topic.
method Introduces latent model with influence and receptivity vectors for each node, estimating topic distributions from observed data.
result The model can estimate topic-based connection structures with theoretical guarantees and outperforms existing methods.
Modeling lead-lag relationship between two text corpora for improved topic modeling.
problem Recognizing the relationship between multiple text corpora for better topic modeling.
method Proposed a jointly dynamic topic model and embedding extension for large-scale text corpus.
result The proposed model can well recognize the lead-lag relationship between two text corpora and improve topic learning.
SCSS detects new events in text streams, overcoming topic modeling shortcomings.
problem Current event detection methods are unsuitable for rapid detection of locally emerging events on massive text streams.
method Alternating optimization between semantic scan and spatial neighborhood discovery.
result SCSS effectively detects real-world disease outbreaks from free-text ED chief complaint data.
Detects changes in topic proportions over time in large text datasets.
problem Unsupervised detection of structural changes in topic distributions over time.
method Specialised temporal topic model with changepoint detection, approximate inference using sample splitting and likelihood ratio statistic.
result Automated detection of changepoints in topic proportions, facilitating interpretable results.
Develops a flexible deep autoencoding topic model with scalable hybrid Bayesian inference.
problem Flexible and interpretable document analysis models.
method DATM with hybrid Bayesian inference, including topic-layer-adaptive stochastic gradient Riemannian MCMC and Weibull variational encoder.
result Demonstrates scalability and efficacy on big corpora in unsupervised and supervised learning tasks.
Proposes TEMN for better POI recommendations.
problem Challenges in capturing user preferences and spatio-temporal POI relationships.
method Integrates topic model and memory network, incorporating geographical module.
result Improves POI recommendation effectiveness by 3.25% and 29.95%.
Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to perform scene recognition and annotation. Recently, a new type of topic model called the Document Neural Autoregressive Distribution Estimator (DocNADE) was proposed and demonstrated state-of-the-art performance for document mod…
Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…
An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. The current practice of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential us…
Inference is an integral part of probabilistic topic models, but is often non-trivial to derive an efficient algorithm for a specific model. It is even much more challenging when we want to find a fast inference algorithm which always yields sparse latent representations of documents. In this article, we introduce a si…
The increasing volume of short texts generated on social media sites, such as Twitter or Facebook, creates a great demand for effective and efficient topic modeling approaches. While latent Dirichlet allocation (LDA) can be applied, it is not optimal due to its weakness in handling short texts with fast-changing topics…
Topic models, such as Latent Dirichlet Allocation (LDA), posit that documents are drawn from admixtures of distributions over words, known as topics. The inference problem of recovering topics from admixtures, is NP-hard. Assuming separability, a strong assumption, [4] gave the first provable algorithm for inference. F…
A new method for scalable inference in deep discrete LVMs.
problem Challenges in scalable inference for deep discrete latent variable models.
method Topic-layer-adaptive stochastic gradient Riemannian MCMC (TLASGR) for DLDA.
result State-of-the-art results on big data sets.
Study uses topic modeling and sentiment analysis to uncover hedge fund performance insights.
problem Hedge fund opacity and limited disclosure make them hard to analyze.
method Applied topic modeling and sentiment analysis to hedge fund documents using DistilBERT and Top2Vec.
result Automated topic modeling and sentiment analysis can predict hedge fund performance.
Variational inference is a very efficient and popular heuristic used in various forms in the context of latent variable models. It's closely related to Expectation Maximization (EM), and is applied when exact EM is computationally infeasible. Despite being immensely popular, current theoretical understanding of the eff…
DAPPER improves scalability of DAP topic model for large corpora.
problem Scaling complex models like DAP to large text corpora.
method Adapted approximate inference techniques for DAP, developing CVI-based EM.
result Significant improvements in model fit and training time without compromising structure.