Algorithm reduces audit costs by identifying best service configurations from biased textual evidence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The study detects deceptive language in business communication using AI.
We contribute the largest publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification. It is collected from 26 fact checking websites in English, paired with textual sources and rich metadata, and labelled for veracity by human expert journalists. We present an in-de…
Paper develops a model for verifying facts in tables without pre-retrieved evidence.
Analysts use vague language in reports to convey useful information about future payoffs.
Study improves keyword forecasting in earnings-call prediction markets.
The paper predicts financial markets using news text and semantic network analysis.
Drug-drug interaction (DDI) is a major cause of morbidity and mortality and a subject of intense scientific interest. Biomedical literature mining can aid DDI research by extracting evidence for large numbers of potential interactions from published literature and clinical databases. Though DDI is investigated in domai…
System segments Form 10-K documents into Item sections for financial analysis.
R package sentometrics analyzes text sentiment for predictions.
We introduce a multimodal visual-textual search refinement method for fashion garments. Existing search engines do not enable intuitive, interactive, refinement of retrieved results based on the properties of a particular product. We propose a method to retrieve similar items, based on a query item image and textual re…
The paper proposes a new recommender system combining ratings and textual reviews.
Based on the assumption that economic complexity is characterised by the interactions of economic agents (who) constantly change their actions and strategies in response to the outcome they mutually create, this paper presents how network models can be used a proxies for the mapping, quantification and analysis of Roma…
In recent years, the interest in Big Data sources has been steadily growing within the Official Statistic community. The Italian National Institute of Statistics (Istat) is currently carrying out several Big Data pilot studies. One of these studies, the ICT Big Data pilot, aims at exploiting massive amounts of textual …
Benchmark assesses forecasting models' ability to use textual context.
Examines parallels between human subjects and texts for causal inference.
Proposes Textual Echo Cancellation to improve speech recognition.
Interpretable text-response modelling for structured outcomes
Textual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embe…
Taxicab correspondence analysis visualizes sparse text data sets.
A new framework predicts stock movements using news sentiment and relational data.
New model uses financial filings to predict bankruptcy, even without MDA sections.
Surveying machine learning methods for economic forecasting.
Paper explores different models for fake news detection.
Since datasets with annotation for novelty at the document and/or word level are not easily available, we present a simulation framework that allows us to create different textual datasets in which we control the way novelty occurs. We also present a benchmark of existing methods for novelty detection in textual data s…
Paper introduces a benchmark for predicting bankruptcy from text data.
MoleculeSTM learns from molecule structures and texts for better drug design.
A vast amount of textual web streams is influenced by events or phenomena emerging in the real world. The social web forms an excellent modern paradigm, where unstructured user generated content is published on a regular basis and in most occasions is freely distributed. The present Ph.D. Thesis deals with the problem …
SAFE detects fake news by analyzing text and image similarities.
An integrated approach is proposed across visual and textual data to both determine and justify a medical diagnosis by a neural network. As deep learning techniques improve, interest grows to apply them in medical applications. To enable a transition to workflows in a medical context that are aided by machine learning,…
Textual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relationship of texts on the same edge to graphically embed text. How…
This study improves stock price prediction using multimodal data.
Predicting and discovering drug-drug interactions (DDIs) is an important problem and has been studied extensively both from medical and machine learning point of view. Almost all of the machine learning approaches have focused on text data or textual representation of the structural data of drugs. We present the first …
Enhances GNNs with text features for better fake news detection.
The performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on …
An important, yet largely unstudied, problem in student data analysis is to detect misconceptions from students' responses to open-response questions. Misconception detection enables instructors to deliver more targeted feedback on the misconceptions exhibited by many students in their class, thus improving the quality…
This study uses NLP to predict stock performance based on analyst reports.
Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is often accompanied by other modalities such as images. We introduce a supervised mult…
Paper proposes a method to improve graph clustering by integrating node textual metadata with node signals in GGMs.
Survey on LLMs for time series analytics across various domains.
MSIN model discovers relevant financial news for time series data.
Paper presents ECL dataset for multi-modal bankruptcy prediction.
InfoSEM infers gene regulatory networks without GT labels, improving performance.
A large body of research into semantic textual similarity has focused on constructing state-of-the-art embeddings using sophisticated modelling, careful choice of learning signals and many clever tricks. By contrast, little attention has been devoted to similarity measures between these embeddings, with cosine similari…
Mathematical reasoning---a core ability within human intelligence---presents some unique challenges as a domain: we do not come to understand and solve mathematical problems primarily on the back of experience and evidence, but on the basis of inferring, learning, and exploiting laws, axioms, and symbol manipulation ru…
Study uses LLMs to optimize VC exit timing after IPO.
Human labeling of data can be very time-consuming and expensive, yet, in many cases it is critical for the success of the learning process. In order to minimize human labeling efforts, we propose a novel active learning solution that does not rely on existing sources of unlabeled data. It uses a small amount of labeled…
Transformer learns CoVaR from financial news, improving systemic risk forecasts.