A scalable topic model for large document collections using MapReduce.
problem Scalability issues in topic modeling for large document collections.
method Correlated Topic Model with variational Expectation-Maximization in MapReduce framework.
result Comparable topic coherences with LDA in MapReduce framework.
Framework uses deep learning to analyze large documents and identify their logical structure.
problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.
Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…
Interactive storytelling connects documents with user constraints.
problem Creating coherent stories from diverse documents.
method Interactive constraints and distance measures based on topic distributions.
result Interactive storytelling outperforms existing methods on multiple datasets.
D-ETM models document topics over time using embeddings and variational inference.
problem Capturing evolving topic patterns in sequential documents.
method Combines D-LDA and word embeddings, using random walk priors and variational inference.
result D-ETM outperforms D-LDA on document completion tasks, learning more diverse and coherent topics.
We present the nested Chinese restaurant process (nCRP), a stochastic process which assigns probability distributions to infinitely-deep, infinitely-branching trees. We show how this stochastic process can be used as a prior distribution in a Bayesian nonparametric model of document collections. Specifically, we presen…
Topic models have proven to be a useful tool for discovering latent structures in document collections. However, most document collections often come as temporal streams and thus several aspects of the latent structure such as the number of topics, the topics' distribution and popularity are time-evolving. Several mode…
We develop a nested hierarchical Dirichlet process (nHDP) for hierarchical topic modeling. The nHDP is a generalization of the nested Chinese restaurant process (nCRP) that allows each word to follow its own path to a topic node according to a document-specific distribution on a shared tree. This alleviates the rigid, …
A new method for topic detection using hierarchical latent tree models.
problem Hierarchical topic detection in document collections.
method Graphical models (HLTMs) with binary variables at different levels representing word co-occurrence patterns and document clusters.
result Captures both general and specific topics at different levels of a hierarchical structure.
The paper develops a decision support system for hierarchical text classification of conference proceedings.
problem Classifying documents with a fixed hierarchical structure of topics.
method Developed a weighted hierarchical similarity function to calculate topic relevance, using entropy of words to estimate weights.
result The weighted hierarchical similarity function improves ranking accuracy compared to other methods.
A dynamic keyword selection model for topic modeling of tweets.
problem Adjusting keywords dynamically to mimic past topics with novelty.
method Generative process selects keywords and documents, trained with variational lower bound and stochastic gradient optimization.
result Keyword-based topic model outperforms a sophisticated baseline model by 67%.
We introduce a Deep Boltzmann Machine model suitable for modeling and extracting latent semantic representations from a large unstructured collection of documents. We overcome the apparent difficulty of training a DBM with judicious parameter tying. This parameter tying enables an efficient pretraining algorithm and a …
We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nodes represent latent topics in text collections and links represent similarity among topics. We descr…
Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…
Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can help us cluster a huge collection into different topics or find a subset of the co…
CCAligned creates a massive web document dataset for cross-lingual research.
problem Identifying comparable documents across different languages.
method Using URL signals to label web documents and mining Common Crawl corpus.
result Release of a dataset with over 392 million URL pairs from 8144 language pairs.
Improved TF-IDF for word relevance in health-care social media documents.
problem Determining word relevance in informal documents.
method Semantic Sensitive TF-IDF (STF-IDF) method.
result Decreased TF-IDF mean error rate by 50% to 13.7%.
An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. The current practice of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential us…
Abstract collects open problems in billiards and symplectic geometry.
problem Open problems in billiards and symplectic geometry.
method Compilation of open problems from discussions.
result Compilation of open problems.
Proposes a new method for evaluating and constructing hierarchical topic models.
problem Evaluation and construction of hierarchical topic models.
method Represent HTM as layers and edges, introduce quality measures, and develop a heterogeneous algorithm.
result The proposed heterogeneous algorithm significantly outperforms baseline approaches.
Brain signals predict user interest in digital content.
problem Finding relevant information from large document collections.
method A brain-information interface using EEG to infer user interest from reading Wikipedia.
result Users' interests can be modeled from brain signals, enabling information recommendation.
Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.
problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.
New method learns more diverse topics from text documents considering paragraph structure.
problem Classic Topic Models ignore word position and use symmetric priors, limiting topic diversity.
method Exploits paragraph structure to distinguish between general and specific topics.
result Shows improved topic diversity and relevance in structured documents.
Develops scalable autoencoder for document networks.
problem Sparse and skewed latent node representations in document relational networks.
method Combines graph Poisson factor analysis with Weibull-based graph inference networks.
result Extracts high-quality hierarchical latent document representations.
CLDA improves topic modeling for large, dynamic datasets.
problem Dynamic topic modeling for large, diverse text streams.
method Data decomposition followed by topic modeling on segments, then clustering.
result Very fast runtime and insight into topic composition over time.
Machine learning automates digitization of historical data.
problem Manual transcription is costly and difficult for large, detailed datasets.
method Apply machine learning techniques for unsupervised layout classification and attention-based neural networks.
result Machine learning can automate the digitization process for historical data.
Paper proposes a method to detect glare in document images.
problem Glare obscures text in document images, hindering recognition.
method Divides document into blocks, collects luminance and histogram features, uses CNN to detect glare.
result High recall and f-score in detecting glare.
We document a mechanism operating in complex adaptive systems leading to dynamical pockets of predictability (``prediction days''), in which agents collectively take predetermined courses of action, transiently decoupled from past history. We demonstrate and test it out-of-sample on synthetic minority and majority game…
We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words.…
Funnelling improves cross-lingual text classification accuracy.
problem Classifying documents in multiple languages more accurately than individual language classifiers.
method A two-tier classification system using posterior probabilities from language-dependent classifiers.
result Funnelling significantly outperforms state-of-the-art baselines in multilingual text classification.
Improved object detection for scientific document images.
problem Current object detectors fail to accurately localize regions in scientific document images.
method Revised R-CNN model with region embedding for fine-grained proposals.
result 17% mAP improvement over standard object detection models.
Paper optimizes summarization of multiple document groups for better distinction.
problem Comparative document summarization to select representative documents from multiple groups.
method Formulated new objective functions based on binary classification and maximum mean discrepancy, using gradient-based optimization.
result Gradient-based optimization outperforms other methods in automatic and crowd-sourced evaluations.
Develops deep NMF models using β-divergences for feature extraction.
problem Inadequate evaluation metrics for deep NMF on diverse datasets.
method Introduces new deep NMF models using Kullback-Leibler divergence.
result Improves feature extraction quality across different types of data.
Machine learning explains text document categorization decisions.
problem Understanding how text documents are categorized by machine learning models.
method Layer-wise relevance propagation (LRP) to trace predictions back to individual words.
result Word-based ML models can be made more comprehensible through LRP.
We document and analyze the empirical facts concerning one of the clearest evidence of speculation in financial trading as observed in the postage collection stamp market. We unravel some of the mechanisms of speculative behavior which emphasize the role of fancy and collective behavior. In our conclusion, we propose a…
Classifies privacy policy segments for better user understanding.
problem Difficulty in understanding privacy policies due to legal jargon.
method Uses machine learning and deep learning techniques to classify privacy policy segments.
result Identifies data practices in privacy policies for better user comprehension.
Proposes supervised topic models for classification and regression from crowds.
problem Ambiguity and noise in annotation tasks, especially with large volumes of documents.
method Develops two supervised topic models and an efficient stochastic variational inference algorithm.
result Empirically demonstrates superior performance over state-of-the-art approaches.
A novel multilayer network approach for text analysis.
problem Clustering documents and finding topics in large collections with metadata and hyperlinks.
method Multilayer Networks and Stochastic Block Models applied to multiple data types.
result Taking into account multiple types of information improves topic and document clustering.
New rational curvature measures for 2-complexes.
problem Measuring curvature in 2-dimensional cell complexes.
method Defined and proved rational curvature invariants.
result Computable rational curvature bounds for 2-complexes.
In this tutorial we explain the inference procedures developed for the sparse Gaussian process (GP) regression and Gaussian process latent variable model (GPLVM). Due to page limit the derivation given in Titsias (2009) and Titsias & Lawrence (2010) is brief, hence getting a full picture of it requires collecting resul…
Workshop on geometry and imagination in Minneapolis.
problem Teaching geometry and imagination concepts.
method Interactive summer workshop led by mathematicians.
result Developed educational materials for teaching geometry and imagination.
Neural model incorporates metadata for better text analysis.
problem Lack of metadata in standard text modeling.
method General neural framework based on topic models.
result Achieves strong performance with metadata.
Deep learning fails in classifying handwritten historical documents, traditional methods perform better.
problem Classifying handwritten historical documents using deep learning methods.
method Traditional and deep learning methods were tested on a large collection of handwritten historical manuscripts.
result Deep learning methods performed poorly, traditional methods performed consistently well.
BERT-based word embeddings improve active learning for text datasets.
problem Efficiently labelling large text datasets for machine learning.
method Evaluation of text representation mechanisms (BERT vs. bag of words) in active learning.
result BERT-based word embeddings significantly improve active learning performance.
ATD detects anomalous topic clusters in high-dimensional text data.
problem Detecting anomalous patterns in high-dimensional discrete data.
method Topic models for detecting anomalous topic clusters.
result Our method can accurately detect anomalous topics and salient features in text data.
Deep neural networks improve face matching across different domains for banking security.
problem Matching facial images from ID documents with selfies for secure transactions.
method A novel deep learning architecture using two CNNs for cross-domain face matching.
result Accuracy rates higher than 93% on the FaceBank dataset.
In this paper, we develop the continuous time dynamic topic model (cDTM). The cDTM is a dynamic topic model that uses Brownian motion to model the latent topics through a sequential collection of documents, where a "topic" is a pattern of word use that we expect to evolve over the course of the collection. We derive an…
INFUSER improves reasoning by self-evolving with a generator and solver that co-learn from unstructured documents.
problem Improving reasoning through self-evolution
method INFUSER uses a generator and solver co-evolving in a document pool to improve reasoning.
result INFUSER outperforms strong self-evolution baselines on Olympiad and SuperGPQA benchmarks.