Proposes a flexible neural recommendation framework for better prediction performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Formulae derived for survival and first passage times in stochastic processes.
Hungarian text processing improved with efficient, accurate NLP pipelines.
Natural Language Processing (NLP) and especially natural language text analysis have seen great advances in recent times. Usage of deep learning in text processing has revolutionized the techniques for text processing and achieved remarkable results. Different deep learning architectures like CNN, LSTM, and very recent…
Enhances medical code predictions for multi-morbidity patients using text classification.
A new text representation model combines CNN and VAE for better semantic extraction.
Detecting depression early from social media texts.
This paper explores NLP techniques for insurance, detailing methods and applications.
Following great success in the image processing field, the idea of adversarial training has been applied to tasks in the natural language processing (NLP) field. One promising approach directly applies adversarial training developed in the image processing field to the input word embedding space instead of the discrete…
Model criticism tool evaluates text coherence and structure in generated long-form text.
AudioPaLM combines text and speech models to improve speech processing and translation.
A novel method extracts topological features from word embeddings for text classification.
This paper analyzes text in financial disclosures to improve financial analysis.
Framework extracts symptoms from EHRs for rapid disease outbreak detection.
Analyzes quantitative finance papers from arXiv using text mining and NLP.
HawkesLLM models text generation with temporal influence, improving semantic alignment under limited memory.
Paper explores reducing precision in SVM for faster text classification.
CATR rationalizes text data to stabilize causal effect estimation.
Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, char…
Many challenges in natural language processing require generating text, including language translation, dialogue generation, and speech recognition. For all of these problems, text generation becomes more difficult as the text becomes longer. Current language models often struggle to keep track of coherence for long pi…
Authorship identification is a process in which the author of a text is identified. Most known literary texts can easily be attributed to a certain author because they are, for example, signed. Yet sometimes we find unfinished pieces of work or a whole bunch of manuscripts with a wide variety of possible authors. In or…
A great variety of text tasks such as topic or spam identification, user profiling, and sentiment analysis can be posed as a supervised learning problem and tackle using a text classifier. A text classifier consists of several subprocesses, some of them are general enough to be applied to any supervised learning proble…
Improved algorithm for misspecified MLMDPs with bounded regret and space/time complexities.
Deep learning uses alphabet frequencies to accurately classify fake news.
LexNLP is an open source Python package focused on natural language processing and machine learning for legal and regulatory text. The package includes functionality to (i) segment documents, (ii) identify key text such as titles and section headings, (iii) extract over eighteen types of structured information like dis…
Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vivid object parts. In…
Improved text-conditioned regression using LLMs and diffusion-based neural processes.
In recent years, there has been an exponential growth in the number of complex documents and texts that require a deeper understanding of machine learning methods to be able to accurately classify texts in many applications. Many machine learning approaches have achieved surpassing results in natural language processin…
BBM models short texts using biterms to improve coherence.
This paper presents an algorithm for pricing perpetual American put options with asset-dependent discounting.
This study uses NLP to detect financial risks from documents.
Trains word embeddings from music and text data to link music contexts.
Textual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relationship of texts on the same edge to graphically embed text. How…
New algorithm learns optimal policy for average reward MDPs with sample complexity matching lower bound.
The paper distinguishes between conditional and marginal processes in language models and discusses conditions for usefulness.
Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural networks, and texts are typically modeled through recurrent neural networks. While the requirement for mo…
RNNs classify text by accumulating evidence on a low-dimensional manifold.
Paper introduces a specialized text classification system for French Open Banking transactions.
The Dirichlet process and its extension, the Pitman-Yor process, are stochastic processes that take probability distributions as a parameter. These processes can be stacked up to form a hierarchical nonparametric Bayesian model. In this article, we present efficient methods for the use of these processes in this hierar…
The paper develops a method to identify LLM-generated text without training.
Given a fixed matrix , where , we study the complexity of sampling from a distribution over all subsets of rows where the probability of a subset is proportional to the squared volume of the parallelepiped spanned by the rows (a.k.a. a determinantal point process). In this task, it is im…
Social media is increasingly used by humans to express their feelings and opinions in the form of short text messages. Detecting sentiments in the text has a wide range of applications including identifying anxiety or depression of individuals and measuring well-being or mood of a community. Sentiments can be expressed…
Bayesian methods improve text annotation quality.
This paper addresses the problem of predicting duration of unplanned power outages, using historical outage records to train a series of neural network predictors. The initial duration prediction is made based on environmental factors, and it is updated based on incoming field reports using natural language processing …
A method for disentangling text representations without supervision.
Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is necessary to exploit datasets containing captioned images, meaning that each image is associated with on…
Paper introduces Latent-CLIP for efficient text-image comparison in latent space.
Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…