Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

3717411,1121,482 · Jun 202019922001200920182026
48 results for supervised topic models

Self-supervised learning excels in topic modeling by being less model-specific.

problem How self-supervised learning discovers useful representations in topic models.
method Applying self-supervised learning objectives to topic model-generated data.
result Self-supervised learning objectives can recover useful posterior information for topic models, outperforming misspecified models.

Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…

2009-12-30abs ↗pdf ↗

Improved neural topic model for semi-supervised learning.

problem Representing textual data in an interpretable manner with limited labeled data.
method Label-Indexed Neural Topic Model (LI-NTM) that combines deep generative models with semi-supervised learning.
result LI-NTM outperforms existing models in document reconstruction and classifier performance.

Supervised topic models improve clinical diagnostics by interpreting cooccurence patterns in count data.

problem Standard supervised Latent Dirichlet Allocation (sLDA) struggles with documents having many more words than labels and lacks effective use of supervised labels.
method Investigates penalized optimization methods to train interpretable sLDA models using recognition networks for faster inference.
result Preliminary results show improved predictions on heldout data for predicting anti-depressant medication success based on patient history.

Proposes supervised topic models for classification and regression from crowds.

problem Ambiguity and noise in annotation tasks, especially with large volumes of documents.
method Develops two supervised topic models and an efficient stochastic variational inference algorithm.
result Empirically demonstrates superior performance over state-of-the-art approaches.

Proposes a new training framework for better interpretable latent models.

problem Lack of balance between generative explanations and label predictions in semi-supervised models.
method Prediction-constrained training objective integrating supervisory signals and stochastic gradient descent.
result Improved prediction quality and interpretable topics compared to previous models.

Predicts EHR features that improve prediction, learning coherent topics.

problem Balancing prediction quality and coherence in EHR topic models.
method Prediction-focused topic model that uses supervisory signal to retain relevant features.
result Prediction-focused topic models learn more coherent topics while maintaining competitive predictions.

Predictive topic models retain only relevant terms for better prediction and topic coherence.

problem Misspecification of topic models leads to poor prediction and topic coherence.
method Uses supervisory signal to select vocabulary terms improving prediction performance.
result Prediction-focused topic models learn more coherent topics while maintaining competitive predictions.

Neural model predicts survival outcomes and reveals feature relationships.

problem Predicting time-to-event outcomes and understanding feature relationships in clinical data.
method Survival and topic modeling combined in a neural network framework.
result Neural survival-supervised topic models achieve competitive accuracy with interpretability.

TS-NMF improves topic models by incorporating user-provided labels.

problem Lack of interpretability in unsupervised topic models.
method Semi-supervised non-negative matrix factorization (TS-NMF) with user-provided labeled examples.
result TS-NMF achieves higher Jaccard similarity scores than unsupervised methods at low supervision rates.

Paper proposes a semi-supervised learning approach for automated topic detection in support and feedback data.

problem Manual analysis of large textual feedback and support data is time-consuming and labor-intensive.
method Combines BERT-based multiclassification with a novel PSHTI model for automated topic and sub-topic inference.
result Automated system for better understanding user voice and engagement patterns in enterprise data.

We present sparse topical coding (STC), a non-probabilistic formulation of topic models for discovering latent representations of large collections of data. Unlike probabilistic topic models, STC relaxes the normalization constraint of admixture proportions and the constraint of defining a normalized likelihood functio…

2012-02-14abs ↗pdf ↗

Max-margin learning is a powerful approach to building classifiers and structured output predictors. Recent work on max-margin supervised topic models has successfully integrated it with Bayesian topic models to discover discriminative latent semantic structures and make accurate predictions for unseen testing data. Ho…

2013-10-10abs ↗pdf ↗

Bayesian Topic Regression models causal inference with text and numerical data.

problem Causal inference using observational text data with both text and numerical confounders.
method Combines supervised Bayesian topic model with Bayesian regression framework, respecting the Frisch-Waugh-Lovell theorem.
result Joint approach recovers ground truth with lower bias than benchmarks, superior prediction results compared to separate approaches.

Proposes a new model to predict antidepressants from EHRs, improving accuracy and interpretability.

problem Improving prediction accuracy and interpretability in topic models for clinical data.
method Prediction-constrained latent Dirichlet allocation framework balancing generative likelihood and prediction accuracy.
result Improved prediction of depression medications from EHRs compared to previous methods.

Develops a flexible deep autoencoding topic model with scalable hybrid Bayesian inference.

problem Flexible and interpretable document analysis models.
method DATM with hybrid Bayesian inference, including topic-layer-adaptive stochastic gradient Riemannian MCMC and Weibull variational encoder.
result Demonstrates scalability and efficacy on big corpora in unsupervised and supervised learning tasks.

Paper extends topic models using neighborhood aggregation for better performance.

problem Extending topic models with pre-trained word embeddings and nonlinear output functions.
method Network view of topic models, neighborhood aggregation algorithm.
result Approach outperforms state-of-the-art supervised Latent Dirichlet Allocation.

Two computational models analyze political topics in social media tweets.

problem Measuring political attention in social media is labor-intensive and restrictive.
method Two computational models: supervised classifier and unsupervised topic model.
result Models provide different benefits: supervised classifier reduces labor, unsupervised model uncovers political and non-political uses.

Supervised topic models simultaneously model the latent topic structure of large collections of documents and a response variable associated with each document. Existing inference methods are based on variational approximation or Monte Carlo sampling, which often suffers from the local minimum defect. Spectral methods …

2016-02-19abs ↗pdf ↗

Better stock market predictions can be made by simpler topic models.

problem Improving stock market prediction accuracy using news articles.
method Empirical and theoretical analysis of supervised and plain LDA models, with a focus on Gibbs sampling and random search.
result Simpler topic models (plain LDA) outperform more complex models (sLDA) in out-of-sample performance.

A new parallel MCMC algorithm improves topic modeling without communication.

problem Quasi-ergodicity problem in topic modeling due to multimodal topic distributions.
method Developed an embarrassingly parallel MCMC algorithm for sLDA by switching topic combination and labeling prediction.
result Out-of-sample prediction performance is comparable to non-parallel sLDA but computation time is significantly reduced.

Hybrid approach combines topic and graph embeddings for legal document clustering.

problem Challenges in classifying legal texts due to domain-specific language and limited labeled data.
method Combines unsupervised topic and graph embeddings with a supervised model.
result Improves clustering quality over text-only or graph-only embeddings.

New learning algorithms for dynamic topic modeling in video analysis.

problem Efficiently processing large volumes of video data for autonomous decisions.
method Two novel learning algorithms based on expectation maximisation and variational Bayes inference.
result Comparison of learning algorithms on real video data.

Corrected moment-based methods improve inference in topic model regression.

problem Inferential difficulties in topic model plug-in workflow for regression.
method Corrected spectral moment methods for LDA, response-weighted word moments.
result Direct identification of regression coefficients without estimating topic shares.

LR-Robot automates SLRs with AI, expert oversight, and multidimensional analysis.

problem Efficient but contextually limited outputs from existing SLR frameworks.
method Human-in-the-loop process, structured knowledge sources, retrieval-augmented generation.
result Empirical demonstration of AI-driven literature synthesis in option pricing.

Proposes a semi-supervised approach to predict user-level sentiments in social media.

problem Detect and analyze sentiment in social media, especially user-level sentiments.
method Semi-supervised approach using a heterogeneous graph built from social networks, incorporating user influences and multiple types of links.
result Predicts user-level sentiments for specific topics more effectively than previous supervised learning approaches.

Two deep learning models improve indoor location prediction from WiFi fingerprints.

problem Indoor location prediction from WiFi fingerprints.
method Convolutional mixture density recurrent neural network and VAE-based semi-supervised learning model.
result Proposed models outperform existing methods in real-world datasets.

Semi-supervised learning is an important and active topic of research in pattern recognition. For classification using linear discriminant analysis specifically, several semi-supervised variants have been proposed. Using any one of these methods is not guaranteed to outperform the supervised classifier which does not t…

2014-11-17abs ↗pdf ↗

Systems rank PubMed abstracts and sentences for RDoC criteria, achieving high mAP and MAA.

problem Lack of RDoC labeled datasets and complex labelling process hinder full use of RDoC framework.
method Attention-based neural topic models, supervised and unsupervised sentence ranking models, BM25, BoW, TF-IDF.
result Best systems achieved 1st rank with 0.86 mAP and 0.58 MAA.

The paper discusses how to improve machine learning models using partial differential equations.

problem Improving the performance and generalization of machine learning models.
method The paper reframes implicit regularization techniques in deep learning as explicit gradient regularization using partial differential equations.
result Explicit regularization using PDEs can lead to better model performance and generalization.

Modeling true and false news diffusion in social networks using homogeneity.

problem Difficulties in distinguishing true from false news in social networks.
method Proposes a Bayesian nonparametric model that incorporates homogeneity of news stories to predict their genuineness.
result Homogeneity values of news stories strongly correlate with their genuineness and content.

We introduce supervised latent Dirichlet allocation (sLDA), a statistical model of labelled documents. The model accommodates a variety of response types. We derive an approximate maximum-likelihood procedure for parameter estimation, which relies on variational methods to handle intractable posterior expectations. Pre…

2010-03-03abs ↗pdf ↗