Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2.5%5.0%7.5%10.0% · Sep 199419922001200920182026
48 results for historical documents

Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font…

2016-01-27abs ↗pdf ↗

Improved handwriting recognition for historical documents with minimal labeled data.

problem Challenges in recognizing historical documents, especially lack of text-line annotations.
method Trained a deep CRNN system on 10% labeled data, augmented with crafted multiscale data, and applied model-based normalization.
result Achieved second best result in ICDAR2017 competition on publicly available READ dataset.

Deep learning fails in classifying handwritten historical documents, traditional methods perform better.

problem Classifying handwritten historical documents using deep learning methods.
method Traditional and deep learning methods were tested on a large collection of handwritten historical manuscripts.
result Deep learning methods performed poorly, traditional methods performed consistently well.

This paper proposes to perform authorship analysis using the Fast Compression Distance (FCD), a similarity measure based on compression with dictionaries directly extracted from the written texts. The FCD computes a similarity between two documents through an effective binary search on the intersection set between the …

2014-02-14abs ↗pdf ↗

Machine learning automates digitization of historical data.

problem Manual transcription is costly and difficult for large, detailed datasets.
method Apply machine learning techniques for unsupervised layout classification and attention-based neural networks.
result Machine learning can automate the digitization process for historical data.

Machine learning identifies types of alterations in historical manuscripts.

problem Understanding and categorizing alterations in historical manuscripts.
method Alteration Latent Dirichlet Allocation (alterLDA) model.
result High performance in recognizing alterations on labelled data, and interesting insights on unlabelled data.

We explore two techniques which use color to make sense of statistical text models. One method uses in-text annotations to illustrate a model's view of particular tokens in particular documents. Another uses a high-level, "words-as-pixels" graphic to display an entire corpus. Together, these methods offer both zoomed-i…

2016-06-20abs ↗pdf ↗

End-to-end solution for recognizing handwritten numerals, avoiding traditional preprocessing steps.

problem Handwritten numeral string recognition with traditional preprocessing steps.
method YoLo-based model for automatic detection and recognition, avoiding heuristic-based preprocessing and segmentation.
result Proposed method reduces complexity and is a feasible end-to-end solution for numeral string recognition.

The study examines how investor protection and past information affect stock returns and interest rates.

problem Empirical regularities related to investor protection and past information in asset pricing models.
method Developed a dynamic asset pricing model with a controlling shareholder and good/bad memory in budget dynamics.
result Good/bad memory of investors on historical market information affects stock returns and interest rates, strengthening investor protection in high ownership concentration.

Framework uses deep learning to analyze large documents and identify their logical structure.

problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.

A new method for document network embedding interprets and generalizes well.

problem Lack of interpretability and generalization to new documents in existing methods.
method Introduces Topic-Word Attention (TWA) and Inductive Document Network Embedding (IDNE) to generate document representations.
result Achieves state-of-the-art performance on various networks and produces meaningful representations.

Study finds Value Granger-causes Size during crisis regimes but not during normal times.

problem Understanding regime-dependent predictive relationships between equity factors.
method Used 35 years of Fama-French data and a Student-t Hidden Markov Model (HMM) to identify crisis regimes.
result Value Granger-causes Size during crisis regimes but not during normal times, validating across multiple historical events.

MarlRank uses multi-agent reinforcement learning to improve document ranking.

problem Neglecting mutual information among documents in ranking models.
method Formulated as a multi-agent Markov Decision Process (MDP), each document predicts relevance considering its own and similar documents features and actions.
result Significant performance gains over state-of-the-art baselines on LETOR benchmark datasets.

New method uses surrogate outcomes and single-record data to improve suicide risk modeling.

problem Lack of historical information in single-record patients hinders modeling rare medical events.
method Hybrid framework combining supervised and unsupervised learning to integrate concurrent and single-record data.
result Single-record data and concurrent diagnoses provide valuable information for improving suicide risk modeling.

Paper develops a new unsupervised scoring function for cross-lingual document alignment.

problem Aligning documents across different languages for NLP tasks.
method Uses cross-lingual sentence embeddings to compute semantic distances and guides document alignment.
result The proposed scoring function outperforms current methods by 7-22% on various language pairs.

IPO Finance Agent extends Finance Agent v2 for SpaceX S-1 filings, improving accuracy and cost-efficiency.

problem Evaluating IPO due diligence tasks with long-form documents.
method Extended task domain, improved agentic harness with contextual retrieval, automated rubric generation.
result Best-performing model reaches 79.8% accuracy, cost-efficient model at 77.2% with 0.05 USD per query.

A scalable topic model for large document collections using MapReduce.

problem Scalability issues in topic modeling for large document collections.
method Correlated Topic Model with variational Expectation-Maximization in MapReduce framework.
result Comparable topic coherences with LDA in MapReduce framework.

A new method for measuring document similarity using hierarchical optimal transport.

problem Inability to measure semantic similarities and scalability issues in past document similarity measures.
method Model documents as distributions over topics, topics as distributions over words, solve optimal transport problem on topics.
result Hierarchical optimal transport provides better interpretability and scalability with comparable performance.

We develop a nested hierarchical Dirichlet process (nHDP) for hierarchical topic modeling. The nHDP is a generalization of the nested Chinese restaurant process (nCRP) that allows each word to follow its own path to a topic node according to a document-specific distribution on a shared tree. This alleviates the rigid, …

2012-10-25abs ↗pdf ↗

Traditional Relational Topic Models provide a way to discover the hidden topics from a document network. Many theoretical and practical tasks, such as dimensional reduction, document clustering, link prediction, benefit from this revealed knowledge. However, existing relational topic models are based on an assumption t…

2015-03-30abs ↗pdf ↗

Study improves document processing in banking with multimodal analytics.

problem Raising operational efficiency in banking through document-intensive processes.
method Comparative analysis of text classifiers and multimodal model (LayoutXLM) on company register extracts.
result Incorporating layout information in a model substantially increases performance.

Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…

2013-09-26abs ↗pdf ↗

This paper improves topic modeling by embedding words and topics together.

problem Topic models struggle with short documents and approximate inference.
method Model each document as a mixture of word embeddings and each topic as a mixture of topic embeddings.
result The method optimizes topic embeddings to minimize semantic differences between words and topics.

Improved bio-surveillance through automated document classification.

problem Tracking infectious diseases across global news alerts.
method Recurrent neural networks, TF-IDF, Naive Bayes, logistic regression.
result 97% recall and 93.3% accuracy in bio-surveillance event classification.

Develops scalable autoencoder for document networks.

problem Sparse and skewed latent node representations in document relational networks.
method Combines graph Poisson factor analysis with Weibull-based graph inference networks.
result Extracts high-quality hierarchical latent document representations.

Paper proposes LAHA to improve XMTC by integrating document content and label correlation.

problem Challenges in tagging documents with most relevant labels from a large label set.
method Hybrid attention deep neural network model (LAHA) that combines multi-label self-attention and adaptive fusion strategies.
result LAHA outperforms state-of-the-art methods, especially on tail labels.

Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.

problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.

Proposes new listwise learning-to-rank models to address rating ties and document relevance.

problem Rating ties and document relevance in existing listwise learning-to-rank models.
method Models ranking as selecting documents from a candidate set based on unique rating levels. Uses a new loss function and adapted RNN model for refining prediction scores.
result Models notably outperform state-of-the-art learning-to-rank models on four public datasets.