This paper introduces a very challenging dataset of historic German documents and evaluates Fully Convolutional Neural Network (FCNN) based methods to locate handwritten annotations of any kind in these documents. The handwritten annotations can appear in form of underlines and text by using various writing instruments…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Research designs first model for Amharic handwritten character recognition.
End-to-end solution for recognizing handwritten numerals, avoiding traditional preprocessing steps.
Deep learning fails in classifying handwritten historical documents, traditional methods perform better.
This paper presents a Convolutional Neural Network (CNN) based page segmentation method for handwritten historical document images. We consider page segmentation as a pixel labeling problem, i.e., each pixel is classified as one of the predefined classes. Traditional methods in this area rely on carefully hand-crafted …
Machine learning automates digitization of historical data.
Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font…
In many real life problems, objects are described by large number of binary features. For instance, documents are characterized by presence or absence of certain keywords; cancer patients are characterized by presence or absence of certain mutations etc. In such cases, grouping together similar objects/profiles based o…
Industry datasets used for text classification are rarely created for that purpose. In most cases, the data and target predictions are a by-product of accumulated historical data, typically fraught with noise, present in both the text-based document, as well as in the targeted labels. In this work, we address the quest…
Historical documents present many challenges for offline handwriting recognition systems, among them, the segmentation and labeling steps. Carefully annotated textlines are needed to train an HTR system. In some scenarios, transcripts are only available at the paragraph level with no text-line information. In this work…
This paper proposes to perform authorship analysis using the Fast Compression Distance (FCD), a similarity measure based on compression with dictionaries directly extracted from the written texts. The FCD computes a similarity between two documents through an effective binary search on the intersection set between the …
KaoKore dataset extracts faces from pre-modern Japanese art for machine learning.
HW2MP-GAN tackles ancient handwritten text recognition.
AD-HOC simplifies high-order derivative calculations in C++.
Machine learning identifies types of alterations in historical manuscripts.
Binarization of digital documents is the task of classifying each pixel in an image of the document as belonging to the background (parchment/paper) or foreground (text/ink). Historical documents are often subjected to degradations, that make the task challenging. In the current work a deep neural network architecture …
Classification Ensemble, which uses the weighed polling of outputs, is the art of combining a set of basic classifiers for generating high-performance, robust and more stable results. This study aims to improve the results of identifying the Persian handwritten letters using Error Correcting Output Coding (ECOC) ensemb…
The paper presents a recognition system for Pashto letters using KNN and ANN.
New handwritten digits dataset for Kannada script.
RNN model predicts handwritten characters from accelerometer and gyroscope data.
A tailored HTR system improves CER to 0.015 for medieval Latin.
In the era of deep learning several unsupervised models have been developed to capture the key features in unlabeled handwritten data. Popular among them is the Restricted Boltzmann Machines RBM. However, due to the novelty in handwritten multidialect data, the RBM may fail to generate an efficient representation. In t…
We explore two techniques which use color to make sense of statistical text models. One method uses in-text annotations to illustrate a model's view of particular tokens in particular documents. Another uses a high-level, "words-as-pixels" graphic to display an entire corpus. Together, these methods offer both zoomed-i…
We propose a simple kernel based nearest neighbor approach for handwritten digit classification. The "distance" here is actually a kernel defining the similarity between two images. We carefully study the effects of different number of neighbors and weight schemes and report the results. With only a few nearest neighbo…
This work attempts to find the most optimal parameter setting of a deep artificial neural network (ANN) for Bengali digit dataset by pre-training it using stacked denoising autoencoder (SDA). Although SDA based recognition is hugely popular in image, speech and language processing related tasks among the researchers, i…
Improved CNN for HCCR with new loss function and ranking method.
The paper approaches the task of handwritten text recognition (HTR) with attentional encoder-decoder networks trained on sequences of characters, rather than words. We experiment on lines of text from popular handwriting datasets and compare different activation functions for the attention mechanism used for aligning i…
EASTER improves OCR efficiency and scalability.
Paper generalizes path signature using fractional calculus for improved machine learning.
Hybrid model learns novel handwritten characters better than neural or symbolic models alone.
EdgeNet improves Arabic numeral classification accuracy to 99.59%.
Word meaning changes over time, depending on linguistic and extra-linguistic factors. Associating a word's correct meaning in its historical context is a central challenge in diachronic research, and is relevant to a range of NLP tasks, including information retrieval and semantic search in historical texts. Bayesian m…
Deep learning improves gender classification from handwriting.
It is proposed a new code for contours of plane images. This code was applied for optical character recognition of printed and handwritten characters. One can apply it to recognition of any visual images.
This text aims to explain general relativity to geometers who have no knowledge about physics. Using handwritten notes by Michel Vaugon, we construct the bases of the theory.
The study examines how investor protection and past information affect stock returns and interest rates.
Framework uses deep learning to analyze large documents and identify their logical structure.
The area of Handwritten Signature Verification has been broadly researched in the last decades, but remains an open research problem. The objective of signature verification systems is to discriminate if a given signature is genuine (produced by the claimed individual), or a forgery (produced by an impostor). This has …
A new method for document network embedding interprets and generalizes well.
A new model CDTM improves text classification by concentrating document topics.
New document embedding method finds more relevant tech docs.
Regina is a software package for studying 3-manifold triangulations and normal surfaces. It includes a graphical user interface and Python bindings, and also supports angle structures, census enumeration, combinatorial recognition of triangulations, and high-level functions such as 3-sphere recognition, unknot recognit…
Improved self-supervised learning for document images.
MarlRank uses multi-agent reinforcement learning to improve document ranking.
Study finds Value Granger-causes Size during crisis regimes but not during normal times.
DocParser parses document structures from renderings like PDFs and scans.
New method uses surrogate outcomes and single-record data to improve suicide risk modeling.
CCAligned creates a massive web document dataset for cross-lingual research.