StonkBERT predicts stock price movements using company text data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study evaluates if LLMs have company-specific biases in financial sentiment analysis.
Anonymizing company names in financial news improves trading performance, contrary to initial expectations.
Model earnings call transcripts for better stock price prediction.
Fraud causes substantial costs and losses for companies and clients in the finance and insurance industries. Examples are fraudulent credit card transactions or fraudulent claims. It has been estimated that roughly percent of the insurance industry's incurred losses and loss adjustment expenses each year stem from…
BioFinBERT analyzes sentiment of biotech press releases and financial text around inflection points.
Text classification systems will help to solve the text clustering problem in the Azerbaijani language. There are some text-classification applications for foreign languages, but we tried to build a newly developed system to solve this problem for the Azerbaijani language. Firstly, we tried to find out potential practi…
Study earnings calls to predict stock price movements, finding them more predictive than traditional data.
Ontology learning is a critical task in industry, dealing with identifying and extracting concepts captured in text data such that these concepts can be used in different tasks, e.g. information retrieval. Ontology learning is non-trivial due to several reasons with limited amount of prior research work that automatica…
FinALBERT predicts stock prices using labelled Stocktwits data.
German FinBERT improves financial text analysis performance.
This study reviews text-based stock market analysis methods.
Combining various data types predicts S&P 500 stock prices with high accuracy.
UniFinEval benchmarks financial models across text, images, and videos.
Stock market prediction is one of the most attractive research topic since the successful prediction on the market's future movement leads to significant profit. Traditional short term stock market predictions are usually based on the analysis of historical market data, such as stock prices, moving averages or daily re…
While ubiquitous, textual sources of information such as company reports, social media posts, etc. are hardly included in prediction algorithms for time series, despite the relevant information they may contain. In this work, openly accessible daily weather reports from France and the United-Kingdom are leveraged to pr…
This paper focuses on a comparative evaluation of the most common and modern methods for text classification, including the recent deep learning strategies and ensemble methods. The study is motivated by a challenging real data problem, characterized by high-dimensional and extremely sparse data, deriving from incoming…
LexNLP is an open source Python package focused on natural language processing and machine learning for legal and regulatory text. The package includes functionality to (i) segment documents, (ii) identify key text such as titles and section headings, (iii) extract over eighteen types of structured information like dis…
For sales and marketing organizations within large enterprises, identifying and understanding new markets, customers and partners is a key challenge. Intel's Sales and Marketing Group (SMG) faces similar challenges while growing in new markets and domains and evolving its existing business. In today's complex technolog…
This paper explores how NLP enhances insurance data analysis.
Silent abandonment reduces contact center efficiency by 5%-15%.
Unified model integrates text and time series for financial forecasting.
Study finds similar companies in Dhaka Stock Exchange using technical data.
Study improves document processing in banking with multimodal analytics.
Zero-Copy Architecture Detects Cross-Company Financial Signals Instantly.
MassMutual uses neural network embeddings from financial news to predict downgrade risk.
Predict stock price movements using financial data and news articles with LLMs.
BERTopic improves financial text analysis with FinTextSim's contextual embeddings.
The model is aimed to discriminate the 'good' and the 'bad' companies in Russian corporate sector based on their financial statements data based on Russian Accounting Standards. The data sample consists of 126 Russian public companies- issuers of Ruble bonds which represent about 36% of total number of corporate bonds …
New machine learning method classifies companies effectively.
As companies increase their efforts in retaining customers, being able to predict accurately ahead of time, whether a customer will churn in the foreseeable future is an extremely powerful tool for any marketing team. The paper describes in depth the application of Deep Learning in the problem of churn prediction. Usin…
Employing profits data of Japanese companies in 2002 and 2003, we confirm that Pareto's law and the Pareto index are derived from the law of detailed balance and Gibrat's law. The last two laws are observed beyond the region where Pareto's law holds. By classifying companies into job categories, we find that companies …
Company2Vec creates embeddings from company websites for fine-grained business analytics.
Neural model learns company embeddings from data and news.
Predicting bankruptcy using financial data and news sentiment.
Deep learning models outperform classical methods in forecasting company fundamentals.
The paper develops a decision support system for hierarchical text classification of conference proceedings.
We first estimate the average growth of a company's annual income and its variance by using both real company data and a numerical model which we already introduced a couple of years ago. Investment strategies expecting for income growth is evaluated based on the numerical model. Our numerical simulation suggests the p…
A key factor in developing high performing machine learning models is the availability of sufficiently large datasets. This work is motivated by applications arising in Software as a Service (SaaS) companies where there exist numerous similar yet disjoint datasets from multiple client companies. To overcome the challen…
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of t…
Paper presents a multi-label topic model for financial texts with high performance and insights into market reactions.
Technological change and innovation are vitally important, especially for high-tech companies. However, factors influencing their future research and development (R&D) trends are both complicated and various, leading it a quite difficult task to make technology tracing for high-tech companies. To this end, in this pape…
We introduce a mean-field type approximation for description of company's income statistics. Utilizing huge company data we show that a discrete version of Langevin equation with additive and multiplicative noises can appropriately describe the time evolution of a company's income fluctuation in statistical sense. The …
Analyzes premium data of Indian non-life insurers, finding GEV distribution best fits Lognormal and GEV extremes.
A pairwise clustering approach is applied to the analysis of the Dow Jones index companies, in order to identify similar temporal behavior of the traded stock prices. To this end, the chaotic map clustering algorithm is used, where a map is associated to each company and the correlation coefficients of the financial ti…
In the context of the current financial crisis, when more companies are facing bankruptcy or insolvency, the paper aims to find methods to identify distressed firms by using financial ratios. The study will focus on identifying a group of Romanian listed companies, for which financial data for the year 2008 were availa…
Paper analyzes deep learning models for credit rating prediction using text and numerical data.
Analyzes how inclusion/exclusion from STOXX Europe 600 Index affects company prices.