Applying traditional collaborative filtering to digital publishing is challenging because user data is very sparse due to the high volume of documents relative to the number of users. Content based approaches, on the other hand, is attractive because textual content is often very informative. In this paper we describe …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This is the second installment of the Financial Bubble Experiment. Here we provide the digital fingerprint of an electronic document in which we identify 7 bubbles in 7 different global assets; for 4 of these assets, we present windows of dates of the most likely ending time of each bubble. We will provide that documen…
This is the third installment of the Financial Bubble Experiment. Here we provide the digital fingerprint of an electronic document in which we identify 27 bubbles in 27 different global assets; for 25 of these assets, we present windows of dates of the most likely ending time of each bubble. We will provide that docum…
The tremendous increase in the amount of available research documents impels researchers to propose topic models to extract the latent semantic themes of a documents collection. However, how to extract the hidden topics of the documents collection has become a crucial task for many topic model applications. Moreover, c…
The digitalization of the legal domain has been ongoing for a couple of years. In that process, the application of different machine learning (ML) techniques is crucial. Tasks such as the classification of legal documents or contract clauses as well as the translation of those are highly relevant. On the other side, di…
Machine learning identifies types of alterations in historical manuscripts.
Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font…
Binarization of digital documents is the task of classifying each pixel in an image of the document as belonging to the background (parchment/paper) or foreground (text/ink). Historical documents are often subjected to degradations, that make the task challenging. In the current work a deep neural network architecture …
Finding relevant information from large document collections such as the World Wide Web is a common task in our daily lives. Estimation of a user's interest or search intention is necessary to recommend and retrieve relevant information from these collections. We introduce a brain-information interface used for recomme…
A new test assesses text similarity between two groups of documents.
Machine learning automates digitization of historical data.
In many real life problems, objects are described by large number of binary features. For instance, documents are characterized by presence or absence of certain keywords; cancer patients are characterized by presence or absence of certain mutations etc. In such cases, grouping together similar objects/profiles based o…
Numeracy is the ability to understand and work with numbers. It is a necessary skill for composing and understanding documents in clinical, scientific, and other technical domains. In this paper, we explore different strategies for modelling numerals with language models, such as memorisation and digit-by-digit composi…
Retrieving indexed documents, not by their topical content but their writing style opens the door for a number of applications in information retrieval (IR). One application is to retrieve textual content of a certain author X, where the queried IR system is provided beforehand with a set of reference texts of X. Autho…
BiLRP explains deep similarity models by decomposing scores into feature contributions.
Authorship verification (AV) is a research subject in the field of digital text forensics that concerns itself with the question, whether two documents have been written by the same person. During the past two decades, an increasing number of proposed AV approaches can be observed. However, a closer look at the respect…
Ensuring the security of transactions is currently one of the major challenges that banking systems deal with. The usage of face for biometric authentication of users is attracting large investments from banks worldwide due to its convenience and acceptability by people, especially in cross-domain scenarios, in which f…
Digital money could reduce germ spread during coronavirus.
Study of digital topology concepts like hyperspaces and function graphs.
The paper highlights issues with fixed point claims in digital images.
Find limiting sets for digital cones and suspensions.
Corrects incorrect assertions about fixed points in digital topology.
Study AFPP of unions of convex digital disks in 2D.
The paper addresses flaws in fixed point assertions for digital images.
Study minimal freezing sets in convex digital disks.
Study on cold and freezing sets in digital images.
Incorrect fixed point assertions in digital topology are discussed.
In this paper, we show how to construct graph theoretical models of n-dimensional continuous objects and manifolds. These models retain topological properties of their continuous counterparts. An LCL collection of n-cells in Euclidean space is introduced and investigated. If an LCL collection of n-cells is a cover of a…
Incorrect fixed point assertions in digital topology are discussed.
Digital trees have approximate fixed point property, and conditions for products are explored.
Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be multi-themed, complex, noisy and cover diverse topics. We describe a framework that c…
Study convexity and AFPP in digital images.
The paper highlights issues in fixed point claims in digital topology.
Han discusses variants of digital covering maps and their equivalences.
Critiques incorrect fixed point assertions in digital topology.
Examines how irreducibility and rigidity affect digital images.
Study restrictions on digitally continuous functions and their effects.
A new model CDTM improves text classification by concentrating document topics.
Fixed point assertions in digital topology are often incorrect or poorly stated.
A new method streamlines digital payment programming using smart contracts.
New framework explains leading digit patterns without probabilistic assumptions.
The DAO Report led to a significant shift of ICO activity to Europe.
Study freezing sets for digital images in a 2D grid.
Document network embedding aims at learning representations for a structured text corpus i.e. when documents are linked to each other. Recent algorithms extend network embedding approaches by incorporating the text content associated with the nodes in their formulations. In most cases, it is hard to interpret the learn…
Improved self-supervised learning for document images.
Pakistan examines digital mergers using traditional competition tools.
Wrist movements can reveal digits, posing security risks.
Generative models create personalized patient health simulations.