Faster, context-aware topic modeling approach unveiled.
problem Slow and heavy preprocessing in topic modeling.
method Conceptualizes topics as independent axes, decomposes embeddings using ICA.
result 4.5x faster than BERTopic on average, with diverse and coherent topics.
The simplicial condition and other stronger conditions that imply it have recently played a central role in developing polynomial time algorithms with provable asymptotic consistency and sample complexity guarantees for topic estimation in separable topic models. Of these algorithms, those that rely solely on the simpl…
A new method selects anchor words for better topic discovery in text corpora.
problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.
Study shows LDA topic models converge at rate n^-1/4 without strict topic separability.
problem Convergence rates of Latent Dirichlet Allocation (LDA) topic models.
method Maximum likelihood estimator, Wasserstein's distance metric, without separability or non-degeneracy assumptions.
result Maximum likelihood estimator converges at rate n^-1/4, optimal in worst case.
Develops efficient algorithm for topic discovery with novel geometric insights.
problem Discovering topics from documents with shared latent factors.
method Geometric insights from normalized word co-occurrence matrix and isotropic random projections.
result Provably efficient algorithm with polynomial computation and sample complexity bounds.
New method identifies quark and gluon jets from collider data.
problem Determine quark and gluon jet distributions from collider data.
method Apply topic modeling to jet distributions, using parton shower and theoretical predictions.
result Determined separate quark and gluon jet distributions and spectra.
A new topic modeling method that minimizes topic simplex volume.
problem Topic modeling efficiency and accuracy.
method Reformulates LDA as minimizing topic simplex volume, uses convex relaxation and ADMM.
result Relaxed problem has same global minimum as original under assumptions.
ETM identifies field-specific keywords in text classification.
problem Unsupervised text classification with field-specific keywords.
method Weighted Lasso penalty and pairwise Kullback-Leibler divergence penalty for topic separation.
result ETM improves topic coherence by 22% and 10% compared to LDA.
CorEx learns topics without assumptions, incorporating human input.
problem Complexity and detailed assumptions in topic modeling.
method Information-theoretic framework, anchor words for minimal human input.
result Topics comparable to LDA with minimal human intervention.
New model predicts user preferences from noisy pairwise comparisons.
problem Predicting user preferences from inconsistent and noisy pairwise comparisons.
method Proposes a Mixed Membership Mallows Models (M4) family and uses statistical connections to topic models.
result Empirically competitive model with polynomial sample complexity guarantees.
The paper investigates topic models, ensuring their statistical identifiability and accuracy.
problem Lack of formal theoretical investigation of topic model identifiability and estimation accuracy.
method Proposes a maximum likelihood estimator (MLE) based on integrated likelihood, introducing new geometric identifiability conditions.
result Introduces weaker conditions for topic model identifiability, allowing a broader investigation.
This tutorial covers BSS methods for separating mixed signals.
problem Separating unknown mixed signals.
method Exposes well-known and recent approaches to BSS.
result Unified presentation of BSS methods.
We present algorithms for topic modeling based on the geometry of cross-document word-frequency patterns. This perspective gains significance under the so called separability condition. This is a condition on existence of novel-words that are unique to each topic. We present a suite of highly efficient algorithms based…
Electrical engineer's AI journey due to deep learning convergence.
problem Separate development of AI and pattern recognition.
method Exploration of AI's historical trajectory and intersection with electrical engineering.
result Convergence of AI and electrical engineering due to deep learning.
Topic models, such as Latent Dirichlet Allocation (LDA), posit that documents are drawn from admixtures of distributions over words, known as topics. The inference problem of recovering topics from admixtures, is NP-hard. Assuming separability, a strong assumption, [4] gave the first provable algorithm for inference. F…
A new neural network separates vocals from music accompaniment.
problem Separating vocals from music accompaniment in recordings.
method Self-attention convolutional neural network (CNN) with densely-connected blocks.
result 19.5% relative improvement in vocals separation.
System detects relevant financial news and predictions from unstructured text.
problem Manual extraction of relevant financial information from news is cumbersome and error-prone.
method Topic modeling with LDA, co-reference resolution, multi-paragraph segmentation, and temporal analysis.
result ROUGE-L values for relevant text and predictions/forecasts were 0.662 and 0.982, respectively.
Bayesian Topic Regression models causal inference with text and numerical data.
problem Causal inference using observational text data with both text and numerical confounders.
method Combines supervised Bayesian topic model with Bayesian regression framework, respecting the Frisch-Waugh-Lovell theorem.
result Joint approach recovers ground truth with lower bias than benchmarks, superior prediction results compared to separate approaches.
The paper shows that a preconditioned SPA is robust to noise without the usual dimension-rank condition.
problem Robustness of preconditioned SPA in separable NMF problems with d>r. method Analysis of preconditioned SPA for separable NMF problems with d>r. result The preconditioned SPA is robust to noise without the dimension-rank condition.
Separation of the sources and analysis of their connectivity have been an important topic in EEG/MEG analysis. To solve this problem in an automatic manner, we propose a two-layer model, in which the sources are conditionally uncorrelated from each other, but not independent; the dependence is caused by the causality i…
The paper tackles error event identification in network logs.
problem Identifying error events from network message logs.
method Transformed the problem into topic discovery in documents using a non-parametric change-point detection algorithm.
result The algorithm identifies error events from message logs efficiently and accurately.
A robust algorithm for non-negative matrix factorization (NMF) is presented in this paper with the purpose of dealing with large-scale data, where the separability assumption is satisfied. In particular, we modify the Linear Programming (LP) algorithm of [9] by introducing a reduced set of constraints for exact NMF. In…
When building large-scale machine learning (ML) programs, such as big topic models or deep neural nets, one usually assumes such tasks can only be attempted with industrial-sized clusters with thousands of nodes, which are out of reach for most practitioners or academic researchers. We consider this challenge in the co…
New framework inscribes maximum volume ellipsoid for structured matrix factorization.
problem Structured matrix factorization with columns in unit simplex.
method Maximum volume inscribed ellipsoid (MVIE) via facet enumeration and convex optimization.
result MVIE framework guarantees exact recovery under certain conditions.
A new convex model tackles noisy separable NMF with provable correctness.
problem Noisy separable NMF with multiple data points near basis vectors.
method Smooth separability assumption, convex model, fast gradient method.
result The convex model provably recovers factors in the presence of noise.
In the mixture models problem it is assumed that there are K distributions θ1,…,θK and one gets to observe a sample from a mixture of these distributions with unknown coefficients. The goal is to associate instances with their generating distributions, or to identify the parameters of the hidden distribu…
Develops platforms to analyze social media data for human behavior and emotions.
problem Understanding human behavior and emotions from social media data.
method Self-structuring incremental machine learning, event detection, natural language processing.
result Captured salient topics and events from social media data, validated against news.
Blind source separation (BSS) is one of the most important and established research topics in signal processing and many algorithms have been proposed based on different statistical properties of the source signals. For second-order statistics (SOS) based methods, canonical correlation analysis (CCA) has been proved to…
We introduce supervised latent Dirichlet allocation (sLDA), a statistical model of labelled documents. The model accommodates a variety of response types. We derive an approximate maximum-likelihood procedure for parameter estimation, which relies on variational methods to handle intractable posterior expectations. Pre…
Paper finds better words for topic models by reranking top words.
problem Top words in topic models are not always representative.
method Reranking words by considering marginal probability over every topic.
result Reranked top words are more representative of topics.
Proposes a new multi-layer model for topic distributions.
problem Leveraging deep structures for learning word distributions of topics.
method A multi-layer generative process on word distributions of topics, where each topic is drawn from a mixture of topics from the layer above.
result Discover interpretable topic hierarchies and improve topic models' accuracy and interpretability.
TBIP uses texts to quantify lawmakers' political positions.
problem Quantifying lawmakers' political positions from speeches, tweets, etc.
method Unsupervised probabilistic topic model analyzing texts.
result TBIP separates lawmakers by party and infers ideal points close to vote-based.
Topic-aware chatbot learns from NMF topic vectors.
problem Improving chatbot relevance based on user topics.
method Combines RNN with NMF for topic learning and attention.
result Chatbot provides more relevant answers based on topic.
Efficiently models correlated topics with topic embeddings.
problem High computational cost and poor scaling in correlated topic modeling.
method Compact topic embeddings and efficient inference in low-dimensional space.
result Handles larger model and data scales without sacrificing performance.
A new model CDTM improves text classification by concentrating document topics.
problem Unsupervised text classification with diverse topic distributions.
method Imposes an exponential entropy penalty on document topic distribution to encourage concentration.
result More coherent topics and concentrated, sparse document-topic distributions.
New metric correlates local topic quality with human judgments.
problem Evaluation of topic models focuses on global metrics, ignoring token-level assignments.
method Proposed a human evaluation task and automated metrics to assess local topic quality.
result Consistency metric correlates best with human judgments of local topic quality.
Proposes HMHP for joint modeling of user-topic interactions.
problem Complex interactions between users, topics and time on social media.
method Hidden Markov Hawkes Process (HMHP) incorporating topical Markov Chains.
result HMHP outperforms state-of-the-art models in generalization and accuracy.
Paper tackles unsupervised learning under latent label shift across domains.
problem Discovering classes from unlabeled data with shifting label distributions.
method Introduces unsupervised learning under Latent Label Shift (LLS), leveraging domain-discriminative models.
result Proves that with domain information, unsupervised classification can improve upon standard methods.
Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…
ETM discovers interpretable topics in large vocabularies.
problem Existing topic models fail with large, heavy-tailed vocabularies.
method Generative model combining topic models and word embeddings with variational inference.
result ETM discovers interpretable topics even with large vocabularies.
Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can help us cluster a huge collection into different topics or find a subset of the co…
Predictive topic models retain only relevant terms for better prediction and topic coherence.
problem Misspecification of topic models leads to poor prediction and topic coherence.
method Uses supervisory signal to select vocabulary terms improving prediction performance.
result Prediction-focused topic models learn more coherent topics while maintaining competitive predictions.
Paper compares NMF and LDA for topic labeling in customer communications.
problem Automatically labeling topics in customer inquiries.
method Uses NMF and LDA for topic mining and labeling.
result Proposes methods for automated topic subject labeling.
New models discover new topics over time in topic modeling.
problem Discovering new topics over time in topic modeling.
method Nonparametric Bayesian models and Hungarian matching algorithm.
result Significantly faster than existing methods, discovering new topics in large datasets.
We introduce Gaussian Process Topic Models (GPTMs), a new family of topic models which can leverage a kernel among documents while extracting correlated topics. GPTMs can be considered a systematic generalization of the Correlated Topic Models (CTMs) using ideas from Gaussian Process (GP) based embedding. Since GPTMs w…
TopicEq model generates equations and text from scientific papers.
problem Communicating ideas in scientific texts using both mathematics and text.
method Joint topic and equation generation model using correlated topic model and RNN.
result Joint model outperforms existing topic and equation models for scientific texts.
New metric R2 improves topic model evaluation.
problem Lack of standard cross-contextual evaluation metrics for topic modeling.
method Introduces R2 as a coefficient of determination for topic models. result Improves topic model evaluation by providing a standard metric.
AOBTM adapts online topic modeling for short app reviews, revealing coherent topics over time.
problem Challenges in inferring latent topics from short, dynamic app reviews over multiple versions.
method Adaptive Online Biterm Topic Model (AOBTM) that addresses sparsity and considers statistical data from previous versions.
result AOBTM finds more coherent topics and outperforms state-of-the-art baselines.