Framework uses deep learning to analyze large documents and identify their logical structure.
problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.
DocParser parses document structures from renderings like PDFs and scans.
problem Parsing complete hierarchical document structures from renderings.
method End-to-end system with novel weak supervision approach.
result Significant improvement in document structure parsing performance.
Deep learning captures semantic structure of large documents.
problem Understanding complex, structured documents like scholarly articles and business reports.
method Deep learning-based document ontology to capture semantic structure and domain-specific concepts.
result The ontology enhances semantic indexing for better understanding by humans and machines.
A new model analyzes document structure and customer shopping patterns.
problem Understanding document structure and customer shopping patterns.
method Variational EM algorithm for multilayer correlated topic modeling.
result MCTM successfully captures document structure and customer shopping patterns.
New language models improve document coherence.
problem Existing language models fail to account for discourse structure.
method Introduced Document-Context Language Models (DCLM) using multi-level recurrent neural networks.
result DCLM models yield better document coherence than word-level models.
New method learns more diverse topics from text documents considering paragraph structure.
problem Classic Topic Models ignore word position and use symmetric priors, limiting topic diversity.
method Exploits paragraph structure to distinguish between general and specific topics.
result Shows improved topic diversity and relevance in structured documents.
System suggests clinical concepts in real-time for faster note creation.
problem Efficiently creating structured clinical notes with minimal keystrokes.
method Contextual autocompletion using shallow neural networks.
result Reduces keystrokes by 67% in real hospital environments.
We introduce a novel latent grouping model for predicting the relevance of a new document to a user. The model assumes a latent group structure for both users and documents. We compared the model against a state-of-the-art method, the User Rating Profile model, where only users have a latent group structure. We estimat…
Improved self-supervised learning for document images.
problem Performance of self-supervised pre-training on document images is poor.
method Proposed context-aware alternatives and a novel multi-modal method.
result Novel method outperforms other self-supervised methods on document image classification.
Traditional Relational Topic Models provide a way to discover the hidden topics from a document network. Many theoretical and practical tasks, such as dimensional reduction, document clustering, link prediction, benefit from this revealed knowledge. However, existing relational topic models are based on an assumption t…
A new method for document network embedding interprets and generalizes well.
problem Lack of interpretability and generalization to new documents in existing methods.
method Introduces Topic-Word Attention (TWA) and Inductive Document Network Embedding (IDNE) to generate document representations.
result Achieves state-of-the-art performance on various networks and produces meaningful representations.
Develops scalable autoencoder for document networks.
problem Sparse and skewed latent node representations in document relational networks.
method Combines graph Poisson factor analysis with Weibull-based graph inference networks.
result Extracts high-quality hierarchical latent document representations.
Proving that next-token prediction makes language models generate coherent long documents.
problem Understanding why language models generate coherent documents despite focusing on next-token prediction.
method Proving the power of next-token prediction in learning longer-range structure using Recurrent Neural Networks (RNN).
result Optimizing next-token prediction in RNNs yields a model that closely approximates the training distribution, even for long-range coherence.
Topic models have proven to be a useful tool for discovering latent structures in document collections. However, most document collections often come as temporal streams and thus several aspects of the latent structure such as the number of topics, the topics' distribution and popularity are time-evolving. Several mode…
Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.
problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.
Paper proposes LAHA to improve XMTC by integrating document content and label correlation.
problem Challenges in tagging documents with most relevant labels from a large label set.
method Hybrid attention deep neural network model (LAHA) that combines multi-label self-attention and adaptive fusion strategies.
result LAHA outperforms state-of-the-art methods, especially on tail labels.
OLALA automates document layout annotation by selecting ambiguous regions for labeling.
problem Efficiently annotating complex document layouts with limited resources.
method Object-Level Active Learning framework that selects ambiguous regions for labeling and uses semi-automatic correction.
result OLALA significantly boosts model performance and improves annotation efficiency.
New CNN model for fast unsupervised document embedding.
problem Efficient unsupervised document embedding with parallelizable architecture.
method Convolutional Neural Network (CNN) for parallel inference and stochastic forward prediction for unsupervised learning.
result Comparable accuracy to state-of-the-art at significantly reduced computational cost.
Latent Dirichlet Allocation models discrete data as a mixture of discrete distributions, using Dirichlet beliefs over the mixture weights. We study a variation of this concept, in which the documents' mixture weight beliefs are replaced with squashed Gaussian distributions. This allows documents to be associated with e…
Paper accelerates K-means clustering for large sparse document data.
problem Efficiently clustering large-scale sparse document data.
method Designs an AFM algorithm leveraging UCs and inverted-index structure.
result Significantly improves clustering speed for large-scale documents.
Improves retrieval accuracy for hierarchical documents, especially for distant matches.
problem Limited expressive power of dual encoder models in hierarchical retrieval.
method Proves feasibility of DEs for HR, introduces pretrain-finetune recipe to improve long-distance retrieval.
result Pretrain-finetune boosts recall on long-distance pairs from 19% to 76%.
LITE models improve query-document relevance with learnable late interactions.
problem Improving query-document relevance with lower latency and storage.
method Proposes learnable late-interaction models (LITE) that use factorized query and document embeddings followed by a learnable scorer.
result Empirically, LITE outperforms previous late-interaction models in re-ranking tasks.
To date, there have been massive Semi-Structured Documents (SSDs) during the evolution of the Internet. These SSDs contain both unstructured features (e.g., plain text) and metadata (e.g., tags). Most previous works focused on modeling the unstructured text, and recently, some other methods have been proposed to model …
D-ETM models document topics over time using embeddings and variational inference.
problem Capturing evolving topic patterns in sequential documents.
method Combines D-LDA and word embeddings, using random walk priors and variational inference.
result D-ETM outperforms D-LDA on document completion tasks, learning more diverse and coherent topics.
A new method for topic detection using hierarchical latent tree models.
problem Hierarchical topic detection in document collections.
method Graphical models (HLTMs) with binary variables at different levels representing word co-occurrence patterns and document clusters.
result Captures both general and specific topics at different levels of a hierarchical structure.
The paper improves CRFs for logical constraints in structured learning.
problem Structured learning with logical constraints in machine learning.
method General extension of CRF for logical constraints.
result Improved performance on a Document Understanding task.
Hybrid approach combines topic and graph embeddings for legal document clustering.
problem Challenges in classifying legal texts due to domain-specific language and limited labeled data.
method Combines unsupervised topic and graph embeddings with a supervised model.
result Improves clustering quality over text-only or graph-only embeddings.
FUNSD dataset tackles noisy scanned forms, offering comprehensive annotations.
problem Extracting and structuring textual content from noisy scanned documents.
method Comprehensive dataset with real, fully annotated forms, including text detection, OCR, layout analysis, and entity linking.
result First publicly available dataset for form understanding, addressing challenges in noisy scanned documents.
Unsupervised scheme ranks sentences in text documents based on semantic importance.
problem Ranking sentences in text documents without labeled data.
method Extracts essential words and phrases, constructs semantic phrase and sentence graphs, applies PageRank, combines scores, and optimizes for topic diversity.
result SSR outperforms individual judges and compares favorably with combined rankings on benchmarks.
Improves document summarization by combining word embeddings and n-grams.
problem Exact word matching fails to measure semantic similarity between sentences.
method Uses deep embedding features and tf-idf features to improve sentence similarity measure; builds an improved sentence similarity graph; employs a submodular objective function; develops a Transformer-based compression model.
result Outperforms tf-idf based approach and achieves state-of-the-art performance on DUC04 dataset.
We present the nested Chinese restaurant process (nCRP), a stochastic process which assigns probability distributions to infinitely-deep, infinitely-branching trees. We show how this stochastic process can be used as a prior distribution in a Bayesian nonparametric model of document collections. Specifically, we presen…
Paper improves matrix factorization for remote sensing and document clustering.
problem Volume minimization for structured matrix factorization.
method Developed a new VolMin algorithm for robust matrix factorization, addressing computational complexity and outlier sensitivity.
result The new algorithm effectively handles volume regularization and iteratively downweights outliers.
The paper develops a decision support system for hierarchical text classification of conference proceedings.
problem Classifying documents with a fixed hierarchical structure of topics.
method Developed a weighted hierarchical similarity function to calculate topic relevance, using entropy of words to estimate weights.
result The weighted hierarchical similarity function improves ranking accuracy compared to other methods.
HC test measures word-frequency similarity for authorship attribution.
problem Identifying the author of a document based on word-frequency patterns.
method Adapting Higher Criticism (HC) to compare word-frequency tables.
result HC identifies characteristic words of the author, unaffected by topic structure.
Graph ConvNet improves classification by leveraging label graph structure.
problem Ignoring label graph structure in multi-class classification leads to suboptimal performance.
method Proposes a GCN-based neural network classifier that incorporates the graph structure of labels.
result The proposed model outperforms baseline methods in terms of graph-theoretic metrics.
Entity-GCN model answers multi-document questions by reasoning across documents.
problem Answering questions based on multiple documents and cross-document relations.
method Graph Convolutional Networks (GCNs) applied to a graph of mentions and their relations.
result Achieves state-of-the-art results on WikiHop dataset.
Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can help us cluster a huge collection into different topics or find a subset of the co…
A new model CDTM improves text classification by concentrating document topics.
problem Unsupervised text classification with diverse topic distributions.
method Imposes an exponential entropy penalty on document topic distribution to encourage concentration.
result More coherent topics and concentrated, sparse document-topic distributions.
New document embedding method finds more relevant tech docs.
problem Computing document similarities for technical content.
method Hybrid approach using graph techniques to extract keyphrases and score sentences.
result Proposed methods outperform baselines up to 27% in NDCG.
Study examines effects of pruning techniques on deep learning models.
problem Understanding the impact of pruning methods on deep learning model structure and dynamics.
method Investigated differences in connectivity and learning dynamics of pruned models using various iterative pruning techniques.
result Emergence of structure in pruned models through magnitude-based unstructured pruning and weight rewinding.
Automated test model creation from semi-structured requirements.
problem Lack of automated solution for test model creation from requirements.
method Machine Learning for semi-structured requirement detection and rule-based translation.
result 86% time savings with no loss of quality.
We introduce a method to learn a mixture of submodular "shells" in a large-margin setting. A submodular shell is an abstract submodular function that can be instantiated with a ground set and a set of parameters to produce a submodular function. A mixture of such shells can then also be so instantiated to produce a mor…
MarlRank uses multi-agent reinforcement learning to improve document ranking.
problem Neglecting mutual information among documents in ranking models.
method Formulated as a multi-agent Markov Decision Process (MDP), each document predicts relevance considering its own and similar documents features and actions.
result Significant performance gains over state-of-the-art baselines on LETOR benchmark datasets.
Generative topic embedding combines local and global patterns for document representation.
problem Representing documents in a continuous space using both local and global patterns.
method Proposes a variational inference model to generate topic embeddings and document representations.
result Performs better than existing methods in document classification tasks.
IllinoisSL is a Java library for learning structured prediction models. It supports structured Support Vector Machines and structured Perceptron. The library consists of a core learning module and several applications, which can be executed from command-lines. Documentation is provided to guide users. In Comparison to …
Much of human knowledge sits in large databases of unstructured text. Leveraging this knowledge requires algorithms that extract and record metadata on unstructured text documents. Assigning topics to documents will enable intelligent search, statistical characterization, and meaningful classification. Latent Dirichlet…
An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. The current practice of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential us…
TopicRNN integrates RNNs and latent topics for better semantic dependency capture.
problem Capturing long-range semantic dependencies in sequential data.
method End-to-end learned RNN with latent topics.
result TopicRNN outperforms existing contextual RNN baselines in word prediction and sentiment analysis.