PCA initializes deep nets for faster, stable document image analysis.
problem Slow and unstable initialization of deep neural networks for document image analysis.
method Turn PCA into an auto-encoder, generating encoder and decoding layers.
result PCA-based initialization leads to faster and more stable network training.
Improved self-supervised learning for document images.
problem Performance of self-supervised pre-training on document images is poor.
method Proposed context-aware alternatives and a novel multi-modal method.
result Novel method outperforms other self-supervised methods on document image classification.
Paper proposes a method to detect glare in document images.
problem Glare obscures text in document images, hindering recognition.
method Divides document into blocks, collects luminance and histogram features, uses CNN to detect glare.
result High recall and f-score in detecting glare.
System identifies high impact research from academic papers.
problem Identifying high impact research amidst a growing volume of papers.
method Combines visual and text classifiers on a large dataset of PDFs and citation counts.
result Improved accuracy in predicting high impact research.
Bayesian non-parametric method clusters high-dimensional binary data.
problem Clustering high-dimensional binary data, especially in complex scenarios.
method Dirichlet Process (DP) mixture model with simulated annealing.
result Outperformed other clustering methods in simulated and real datasets.
OLALA automates document layout annotation by selecting ambiguous regions for labeling.
problem Efficiently annotating complex document layouts with limited resources.
method Object-Level Active Learning framework that selects ambiguous regions for labeling and uses semi-automatic correction.
result OLALA significantly boosts model performance and improves annotation efficiency.
Simple CNN achieves page segmentation for historical documents.
problem Page segmentation of handwritten historical document images.
method Proposes a simple CNN architecture trained on raw image pixels for pixel labeling.
result Simple CNN achieves competitive results compared to other deep architectures.
Improved object detection for scientific document images.
problem Current object detectors fail to accurately localize regions in scientific document images.
method Revised R-CNN model with region embedding for fine-grained proposals.
result 17% mAP improvement over standard object detection models.
Gaussian Process upsampling boosts OCR accuracy from low-res images.
problem Low-quality and downsampled image data hinders OCR accuracy.
method Gaussian Process upsampling model for improving OCR on low-resolution documents.
result Upsampling improves OCR accuracy on low-resolution images.
We develop methods to cluster financial time series using distances between dependent random variables.
problem Inaccurate covariance matrices and difficulty in estimating them from empirical data.
method We propose a new approach to clustering financial time series by using distances between cross-dependent random processes.
result Our method is statistically consistent and can be applied to a broader range of financial analyses.
A new model analyzes document structure and customer shopping patterns.
problem Understanding document structure and customer shopping patterns.
method Variational EM algorithm for multilayer correlated topic modeling.
result MCTM successfully captures document structure and customer shopping patterns.
Bi-LSTM classifies legal documents for Brazil's supreme court.
problem Clogging of Brazil's supreme court due to high volume of lawsuit cases.
method Used a Bidirectional Long Short-Term Memory (Bi-LSTM) network.
result Successfully classified legal documents for efficient case allocation.
FAKTA automates fact checking across media sources.
problem Automating fact checking across diverse media sources.
method Unified framework integrating document retrieval, stance detection, evidence extraction, and linguistic analysis.
result FAKTA predicts factuality and provides evidence for claims.
Active learning reduces font classification data needs for historical documents.
problem Identifying fonts in historical documents for OCR accuracy.
method Active learning strategy using image features and bag-of-word representation.
result Combination of uncertainty and diversity sampling achieves 89% accuracy with 17% labeled data.
Framework uses deep learning to analyze large documents and identify their logical structure.
problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.
New ONMF model minimizes KL divergence for better sparse data modeling.
problem Clustering and data modeling with sparse vectors.
method Developed KL-ONMF algorithm based on alternating optimization.
result KL-ONMF outperforms Frobenius-norm ONMF for document classification and hyperspectral image unmixing.
Develops deep NMF models using β-divergences for feature extraction.
problem Inadequate evaluation metrics for deep NMF on diverse datasets.
method Introduces new deep NMF models using Kullback-Leibler divergence.
result Improves feature extraction quality across different types of data.
Generative model predicts user interest for document recommendations.
problem Improving document recommendation systems.
method Learned document representations and Gaussian mixture model of user interest.
result Generative model outperforms traditional LSA in predictive performance.
Paper benchmarks Bengali language classification tasks using MConv-LSTM network.
problem Lack of computational resources for NLP tasks in under-resourced languages like Bengali.
method Built three datasets, BengFastText word embeddings, and MConv-LSTM network for hate speech detection, document classification, and sentiment analysis.
result BengFastText yields up to 92.30%, 82.25%, and 90.45% F1-scores in document classification, sentiment analysis, and hate speech detection respectively.
This paper improves topic modeling by embedding words and topics together.
problem Topic models struggle with short documents and approximate inference.
method Model each document as a mixture of word embeddings and each topic as a mixture of topic embeddings.
result The method optimizes topic embeddings to minimize semantic differences between words and topics.
Automates U.S. visa petition document classification and RFE response generation.
problem Manual effort in organizing visa petition documents and responding to RFEs.
method Ensemble of image and text classifiers for document categorization and text classifier for RFE evidence identification.
result Achieves considerable accuracy in automated responses while reducing processing time.
SWESA learns word embeddings with document labels for sentiment analysis.
problem Sentiment analysis using limited text data.
method SWESA uses supervised learning to optimize word embeddings and classifier performance.
result SWESA outperforms existing methods in sentiment analysis.
Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.
problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.
Develops scalable autoencoder for document networks.
problem Sparse and skewed latent node representations in document relational networks.
method Combines graph Poisson factor analysis with Weibull-based graph inference networks.
result Extracts high-quality hierarchical latent document representations.
Paper tackles handwritten annotation recognition in historic documents using FCNN.
problem Recognizing handwritten annotations in challenging historic German documents.
method End-to-end semantic segmentation using Fully Convolutional Neural Networks (FCNN).
result Best model achieves 95.6% IoU score on test documents.
Hybrid approach combines topic and graph embeddings for legal document clustering.
problem Challenges in classifying legal texts due to domain-specific language and limited labeled data.
method Combines unsupervised topic and graph embeddings with a supervised model.
result Improves clustering quality over text-only or graph-only embeddings.
Paper improves matrix factorization for remote sensing and document clustering.
problem Volume minimization for structured matrix factorization.
method Developed a new VolMin algorithm for robust matrix factorization, addressing computational complexity and outlier sensitivity.
result The new algorithm effectively handles volume regularization and iteratively downweights outliers.
CNNs show sensitivity to low-frequency signals due to image frequency distribution.
problem Understanding why CNNs are sensitive to low-frequency signals.
method Theoretical analysis of CNN representations in frequency space.
result CNNs sensitivity to low-frequency signals is due to the frequency distribution of natural images.
A new neural topic model using optimal transport improves document representation and topic coherence.
problem Challenges in achieving good document representation and coherent/diverse topics in existing NTMs.
method Proposes a neural topic model via optimal transport, learning topic distribution by minimising OT distance to document word distributions.
result Significantly outperforms state-of-the-art NTMs on discovering coherent and diverse topics.
A deep neural network improves document binarization accuracy.
problem Binarizing digital documents with historical degradations.
method Combines FCN with primal-dual network for end-to-end training.
result Achieves state-of-the-art binarization on four out of seven datasets.
Deep learning fails in classifying handwritten historical documents, traditional methods perform better.
problem Classifying handwritten historical documents using deep learning methods.
method Traditional and deep learning methods were tested on a large collection of handwritten historical manuscripts.
result Deep learning methods performed poorly, traditional methods performed consistently well.
Study uses topic modeling and sentiment analysis to uncover hedge fund performance insights.
problem Hedge fund opacity and limited disclosure make them hard to analyze.
method Applied topic modeling and sentiment analysis to hedge fund documents using DistilBERT and Top2Vec.
result Automated topic modeling and sentiment analysis can predict hedge fund performance.
Transform learning improves K-means clustering for document analysis.
problem Improving K-means clustering for document analysis.
method Embedding K-means clustering loss into transform learning framework and solving jointly using ADMM.
result Improves over state-of-the-art in document clustering.
A new method for measuring document similarity using hierarchical optimal transport.
problem Inability to measure semantic similarities and scalability issues in past document similarity measures.
method Model documents as distributions over topics, topics as distributions over words, solve optimal transport problem on topics.
result Hierarchical optimal transport provides better interpretability and scalability with comparable performance.
Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to perform scene recognition and annotation. Recently, a new type of topic model called the Document Neural Autoregressive Distribution Estimator (DocNADE) was proposed and demonstrated state-of-the-art performance for document mod…
FUNSD dataset tackles noisy scanned forms, offering comprehensive annotations.
problem Extracting and structuring textual content from noisy scanned documents.
method Comprehensive dataset with real, fully annotated forms, including text detection, OCR, layout analysis, and entity linking.
result First publicly available dataset for form understanding, addressing challenges in noisy scanned documents.
This is the second installment of the Financial Bubble Experiment. Here we provide the digital fingerprint of an electronic document in which we identify 7 bubbles in 7 different global assets; for 4 of these assets, we present windows of dates of the most likely ending time of each bubble. We will provide that documen…
Improved bio-surveillance through automated document classification.
problem Tracking infectious diseases across global news alerts.
method Recurrent neural networks, TF-IDF, Naive Bayes, logistic regression.
result 97% recall and 93.3% accuracy in bio-surveillance event classification.
Scoping review of EO-ML methods for causal inference in poverty geography.
problem Lack of thorough documentation and best practices for EO-ML methods in causal analysis.
method Comprehensive scoping review cataloging five principal approaches.
result Detailed protocol for integrating EO data into causal analysis.
Study compares methods for improving document retrieval accuracy.
problem Improving document retrieval accuracy from large corpora.
method Comparison of query expansion, topic models, and active learning.
result Active learning outperforms keyword lists in most settings.
Quantum LTA improves document analysis performance.
problem Improving latent topic analysis in quantum information retrieval.
method Proposed a quantum-motivated LTA method combining geometry and probability.
result Quantum-motivated LTA outperforms LSA on three datasets.
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
Study uses LLMs to generate investor briefs from company reports and SEC filings.
problem Improving data analysis for individual investors.
method Preprocessed data, used gpt-4o model in RAG regime, evaluated by investors.
result LLMs can generate useful investor briefs from company reports and SEC filings.
Plud system reduces labeling time and produces accurate models for uncategorized images.
problem Time-consuming manual labeling of uncategorized images.
method Iterative semi-supervised workflow combining clustering and classification.
result Reduces labeling time and produces accurate models.
A novel multilayer network approach for text analysis.
problem Clustering documents and finding topics in large collections with metadata and hyperlinks.
method Multilayer Networks and Stochastic Block Models applied to multiple data types.
result Taking into account multiple types of information improves topic and document clustering.
New method predicts ICU stay for pancreatitis patients.
problem Predicting ICU stay for pancreatitis patients.
method Survival-supervised topic modeling with elastic-net regularized Cox model and anchor words.
result Our method is as accurate as best baselines but more interpretable.
A new geometrically-motivated algorithm for nonnegative matrix factorization is developed and applied to the discovery of latent "topics" for text and image "document" corpora. The algorithm is based on robustly finding and clustering extreme points of empirical cross-document word-frequencies that correspond to novel …
Transforms data into separable subspaces for clustering.
problem Data is not always separable into subspaces.
method Embeds subspace clustering techniques into transform learning.
result Improves upon state-of-the-art clustering techniques.