Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Nov 199319922001200920182026
48 results for topic distributions

We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words.…

2012-07-11abs ↗pdf ↗

Proposes a new multi-layer model for topic distributions.

problem Leveraging deep structures for learning word distributions of topics.
method A multi-layer generative process on word distributions of topics, where each topic is drawn from a mixture of topics from the layer above.
result Discover interpretable topic hierarchies and improve topic models' accuracy and interpretability.

LLMs encode latent topic distributions, suggesting Bayesian inference.

problem Capturing topic structure from large language models.
method Connecting LLM optimization to implicit Bayesian inference and de Finetti's theorem.
result LLMs recover latent topic distributions, matching LDA-generated topics.

A new parallel MCMC algorithm improves topic modeling without communication.

problem Quasi-ergodicity problem in topic modeling due to multimodal topic distributions.
method Developed an embarrassingly parallel MCMC algorithm for sLDA by switching topic combination and labeling prediction.
result Out-of-sample prediction performance is comparable to non-parallel sLDA but computation time is significantly reduced.

We describe our language-independent unsupervised word sense induction system. This system only uses topic features to cluster different word senses in their global context topic space. Using unlabeled data, this system trains a latent Dirichlet allocation (LDA) topic model then uses it to infer the topics distribution…

2013-02-28abs ↗pdf ↗

A new topic model uses word embeddings on a sphere for better topic coherence.

problem Traditional topic models ignore semantic word correlations.
method Proposes von Mises-Fisher distribution for word density on a unit sphere, using Hierarchical Dirichlet Process and Stochastic Variational Inference.
result The model outperforms existing methods in topic coherence.

Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.

problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.

A novel method learns word embeddings and topics using Wasserstein distance.

problem Learning word embeddings and topics from text data.
method Distilled Wasserstein learning framework for joint word embedding and topic modeling.
result Superior performance on disease network construction, mortality prediction, and procedure recommendation.

A new neural topic model using optimal transport improves document representation and topic coherence.

problem Challenges in achieving good document representation and coherent/diverse topics in existing NTMs.
method Proposes a neural topic model via optimal transport, learning topic distribution by minimising OT distance to document word distributions.
result Significantly outperforms state-of-the-art NTMs on discovering coherent and diverse topics.

Paper proposes a method to improve language model performance on unknown distributions.

problem Language models trained on diverse data can perform poorly on unseen distributions.
method Distributionally robust optimization (DRO) to minimize worst-case performance over a mixture of potential test distributions.
result Topic CVaR approach reduces perplexity by 5.5 points compared to standard maximum likelihood.

We describe a new method for visualizing topics, the distributions over terms that are automatically extracted from large text corpora using latent variable models. Our method finds significant nn-grams related to a topic, which are then used to help understand and interpret the underlying distribution. Compared with …

2009-07-06abs ↗pdf ↗

Efficiently trains HDP topic models on large datasets using a sparse data-parallel sampler.

problem Scaling non-parametric topic models to large datasets.
method Data-parallel training with a doubly sparse sampler for HDP topic models.
result Trains HDP topic models on a 8m document, 768m token PubMed corpus in under 4 days.

Develops a flexible deep autoencoding topic model with scalable hybrid Bayesian inference.

problem Flexible and interpretable document analysis models.
method DATM with hybrid Bayesian inference, including topic-layer-adaptive stochastic gradient Riemannian MCMC and Weibull variational encoder.
result Demonstrates scalability and efficacy on big corpora in unsupervised and supervised learning tasks.

Neural model predicts survival outcomes and reveals feature relationships.

problem Predicting time-to-event outcomes and understanding feature relationships in clinical data.
method Survival and topic modeling combined in a neural network framework.
result Neural survival-supervised topic models achieve competitive accuracy with interpretability.

Paper proposes a new model to infer topic-based connection structures from noisy adjacency matrices.

problem Traditional community detection assumes static connection structures, but real-world connections vary based on topic.
method Introduces latent model with influence and receptivity vectors for each node, estimating topic distributions from observed data.
result The model can estimate topic-based connection structures with theoretical guarantees and outperforms existing methods.

The logistic normal distribution has recently been adapted via the transformation of multivariate Gaus- sian variables to model the topical distribution of documents in the presence of correlations among topics. In this paper, we propose a probit normal alternative approach to modelling correlated topical structures. O…

2014-10-03abs ↗pdf ↗

This paper improves topic model estimation for sparse distributions and applies it to Wasserstein distances.

problem Estimating sparse topic distributions in topic models with high-dimensional data.
method MLE for topic weights when AA is known, plug-in estimator for unknown AA.
result MLE can be exactly sparse and contain true zero pattern of topic weights.

A single, stationary topic model such as latent Dirichlet allocation is inappropriate for modeling corpora that span long time periods, as the popularity of topics is likely to change over time. A number of models that incorporate time have been proposed, but in general they either exhibit limited forms of temporal var…

2012-08-22abs ↗pdf ↗

Scalable algorithm for extracting topic hierarchies from large text corpora.

problem Efficient inference for hierarchical topic models on large datasets.
method Partially collapsed Gibbs sampling (PCGS) algorithm combined with efficient distributed implementation.
result 111 times more efficient than previous implementation for hLDA.

We examine the problem of learning a probabilistic model for melody directly from musical sequences belonging to the same genre. This is a challenging task as one needs to capture not only the rich temporal structure evident in music, but also the complex statistical dependencies among different music components. To ad…

2012-06-27abs ↗pdf ↗

In latent Dirichlet allocation (LDA), topics are multinomial distributions over the entire vocabulary. However, the vocabulary usually contains many words that are not relevant in forming the topics. We adopt a variable selection method widely used in statistical modeling as a dimension reduction tool and combine it wi…

2012-05-04abs ↗pdf ↗

Enhances topic models to better handle polysemous words.

problem Lack of polysemy handling in Gaussian latent Dirichlet allocation.
method Introduces a hierarchical structure to capture polysemy in Gaussian latent Dirichlet allocation.
result Significantly improves polysemy detection and provides more parsimonious topic representations.

Detects changes in topic proportions over time in large text datasets.

problem Unsupervised detection of structural changes in topic distributions over time.
method Specialised temporal topic model with changepoint detection, approximate inference using sample splitting and likelihood ratio statistic.
result Automated detection of changepoints in topic proportions, facilitating interpretable results.

We present an LDA approach to entity disambiguation. Each topic is associated with a Wikipedia article and topics generate either content words or entity mentions. Training such models is challenging because of the topic and vocabulary size, both in the millions. We tackle these problems using a novel distributed infer…

2013-09-02abs ↗pdf ↗

Develops efficient algorithm for topic discovery with novel geometric insights.

problem Discovering topics from documents with shared latent factors.
method Geometric insights from normalized word co-occurrence matrix and isotropic random projections.
result Provably efficient algorithm with polynomial computation and sample complexity bounds.

nnLDA combines neural and probabilistic methods for better topic modeling with side information.

problem Lack of integration of auxiliary information in traditional topic models.
method nnLDA integrates side information through a neural prior mechanism, optimizing both neural and probabilistic components.
result nnLDA outperforms traditional models in topic coherence, perplexity, and classification.

Recent advances in topic models have explored complicated structured distributions to represent topic correlation. For example, the pachinko allocation model (PAM) captures arbitrary, nested, and possibly sparse correlations between topics using a directed acyclic graph (DAG). While PAM provides more flexibility and gr…

2012-06-20abs ↗pdf ↗

A new nonparametric model discovers hidden topics in document networks without knowing the number of topics in advance.

problem Discovering hidden topics in document networks without knowing the number of topics in advance.
method Proposes a nonparametric relational topic model using dependent gamma processes to represent topic interests of documents and their network structure.
result Can discover hidden topics and their number simultaneously using a posterior inference algorithm.

New tools for estimating and inferring Wasserstein distance in topic models.

problem Estimating and inferring the Wasserstein distance between mixing measures in topic models.
method New canonical interpretation and tools for inference on Wasserstein distance in topic models.
result First minimax lower bounds and fully data-driven inferential tools for the Wasserstein distance in topic models.