Model estimates non-reported GHG emissions for companies using machine learning.
problem Incomplete GHG emissions reporting by companies.
method Interpretable machine learning model tailored for non-reporting companies.
result Model accurately estimates emissions for diverse company groups.
Study clusters Kenyan medical insurance companies based on financial performance and reporting consistency.
problem Identifying financial health and reporting consistency in Kenyan medical insurance companies.
method Advanced clustering techniques (KMeans, DTW) on financial ratios and time series data.
result Four distinct clusters identified, each representing different financial performance and reporting consistency combinations.
FinBERT-XRC model assesses financial report risk, offering transparent explanations.
problem Assessing post-event return volatility risk in financial reports.
method Deep-learning model FinBERT-XRC with explainability at word, sentence, and corpus levels.
result FinBERT-XRC outperforms state-of-the-art models in predictive accuracy.
Study uses LLMs to generate investor briefs from company reports and SEC filings.
problem Improving data analysis for individual investors.
method Preprocessed data, used gpt-4o model in RAG regime, evaluated by investors.
result LLMs can generate useful investor briefs from company reports and SEC filings.
Over the last 23 years, the U.S. Securities and Exchange Commission has required over 34,000 companies to file over 165,000 annual reports. These reports, the so-called "Form 10-Ks," contain a characterization of a company's financial performance and its risks, including the regulatory environment in which a company op…
Develops Merton's model for private companies using DDM.
problem Lack of observable asset values for private companies.
method Uses dividend discount model (DDM) to develop structural model.
result Obtains closed-form formulas for equity and liability values, default probability.
Weak predictability of stock price movement 2 days after annual report disclosure.
problem Predicting stock price movement after annual report disclosure.
method Used various models including decision tree, logistic regression, random forest, neural network, prototypical networks; used financial indicators from EastMoney.
result Maximum accuracy and precision of stock price movement prediction is around 59.6% and 0.56 respectively, with random forest performing best.
Study tests UK FTSE-listed companies' financial data for Benford's Law conformity.
problem Ensuring the fairness of public revenue collection and reducing tax avoidance risks.
method Utilised pre-tax income and total assets data from 567 FTSE companies, tested for Benford's Laws conformity using χ2 and MAD tests. result MAD test rejects Benford's Laws conformity, suggesting potential issues with reported financial data.
Study finds companies react negatively to material cybersecurity incident disclosures.
problem Understanding market reactions to cybersecurity incidents.
method Examined daily stock price movements of companies disclosing material cybersecurity incidents.
result Companies tend to experience negative price reactions after disclosing material cybersecurity incidents.
Large language models learn company embeddings from SEC filings.
problem Lack of a rigorous definition of company similarity.
method Pre-trained and finetuned large language models (LLMs) to learn embeddings from SEC filings.
result LLMs can reproduce GICS classifications and indicate similar financial performance.
System constructs public competitor graph from financial reports.
problem Time-consuming and expert-laden manual extraction of corporate relationships.
method Financial report processing to generate reliable knowledge graph of corporate relationships.
result More than 83% of S\&P 500 companies' competition relationships retrieved.
The paper analyzes how news sentiment of companies can affect market movements.
problem Understanding how news sentiment impacts market performance and volatility.
method Applied NLP techniques to analyze news sentiment of 87 companies over 7 years.
result Strong media sentiment towards one company can indicate significant changes in sentiment towards related companies.
The study visualizes Spanish fish and meat processing companies using financial, environmental, and social ratios.
problem Mapping financial, environmental, and social performance of Spanish processing companies.
method Used compositional data and principal-component analysis biplot for statistical analysis.
result Identified clusters of companies with similar financial, environmental, and social performance.
On a periodic basis, publicly traded companies are required to report fundamentals: financial data such as revenue, operating income, debt, among others. These data points provide some insight into the financial health of a company. Academic research has identified some factors, i.e. computed features of the reported d…
AI uses KGs to assess economic impact of selective lockdowns on Italian companies.
problem Impact of selective lockdowns on Italian companies' economic stability.
method Automated Reasoning and Knowledge Graphs to analyze company networks.
result Identifies strategic companies at risk of takeover during lockdowns.
RAG-IT automates financial analysis using LLMs and specialized datasets.
problem Manual financial analysis is time-consuming and requires expertise.
method Retrieval-Augmented Instruction Tuning (RAG-IT) fine-tunes an LLM for financial tasks.
result RAG-IT improves financial report generation performance compared to commercial systems.
System segments Form 10-K documents into Item sections for financial analysis.
problem Segmenting Form 10-K documents into Item sections for efficient financial analysis.
method Developed an automatic Form 10-K Itemization system using NLP techniques.
result System achieves a retrieval rate of 93% for segmenting Item sections.
StonkBERT predicts stock price movements using company text data.
problem Can language models predict medium-run stock price movements?
method Fine-tuning transformer-based language models (BERT) on company text data (news articles, blogs, annual reports) for stock price performance classification.
result StonkBERT shows substantial improvement in predictive accuracy compared to traditional models, with news articles providing the best results.
Study shows environmental spending positively impacts company profitability.
problem Impact of environmental spending on company profitability.
method Panel data regression analysis using E-Views.
result Environmental spending positively impacts profitability metrics.
Study earnings calls to predict stock price movements, finding them more predictive than traditional data.
problem Improving investment decisions by analyzing earnings calls for stock price predictions.
method Graph Neural Network based approach to process and analyze earnings call transcripts.
result Earnings call transcripts are more predictive of stock price movements than traditional hard data.
TinyXRA assesses financial risks from 10-K reports using a lightweight transformer model.
problem Comprehensive risk assessment from financial reports, distinguishing between upside and downside risk.
method Lightweight transformer model with dynamic attention, incorporating skewness, kurtosis, and Sortino ratio.
result State-of-the-art predictive accuracy and transparent risk assessments.
This paper introduces a non-parametric framework to statistically examine how news events, such as company or macroeconomic announcements, contribute to the pre- and post-event jump dynamics of stock prices under the intraday seasonality of the news and jumps. We demonstrate our framework, which has several advantages …
We report the proof that the extension of Gibrat's law in the middle scale region is unique and the probability distribution function (pdf) is also uniquely derived from the extended Gibrat's law and the law of detailed balance. In the proof, two approximations are employed. The pdf of growth rate is described as tent-…
The paper proposes an original methodology for constructing quantitative statistical models based on multidimensional distribution functions constructed on the basis of the insurance companies' data on inshurance policies (including policies with deductible) and claims incurred. Real data of some Russian insurance comp…
New method decomposes profits and losses continuously, avoiding discrete reporting issues.
problem Analyzing profits and losses at discrete dates ignores detailed paths.
method Constructs a large class of continuous-time decompositions using extended Itô's formula.
result Identifies a preferred decomposition from exactness, symmetry, and normalization axioms.
Paper uses LLMs to analyze annual reports for stock investment, improving efficiency.
problem Manual analysis of annual reports is time-consuming and requires expertise.
method Leverages Large Language Models to extract and analyze annual reports.
result Machine Learning model trained on LLM outputs outperforms S&P500 returns.
Study examines the impact of employment benefit costs on firm profitability.
problem Impact of employment benefit costs on firm profitability.
method Panel data regression analysis using E-Views.
result There is a significant positive relationship between employment benefit costs and firm profitability.
Paper analyzes strategic underreporting in competitive insurance markets.
problem Strategic underreporting by insureds in competitive insurance markets.
method Develops a dynamic insurance market model with two competing companies and a continuum of insureds, examines the interaction between strategic underreporting and competitive pricing under a Bonus-Malus System framework.
result Establishes the existence and uniqueness of the insureds' optimal reporting barrier and its dependence on BMS premiums; proves the existence of Nash equilibrium premium strategies.
Model predicts ESG ratings from news articles using multivariate timeseries analysis.
problem Lack of accurate and automated methods for ESG ratings prediction.
method Multivariate timeseries analysis combined with deep learning.
result Model outperforms state-of-the-art methods in predicting ESG ratings.
New method uses LLMs to extract financial insights from Q&A sections of reports.
problem Scalability and accuracy issues in extracting valuable insights from financial report Q&A sections.
method Combines retrieval-augmented generation technique with metadata.
result Empirically demonstrates superior performance of the proposed method.
CNN predicts stock fluctuations using company news headlines.
problem Predicting next-day stock fluctuations based on company-specific news.
method Convolutional Neural Network (CNN) with reduced filter dimensions and multiple hidden layers. Fine-tuned word embeddings and various filter widths.
result 61.7% classification accuracy achieved using pre-learned embeddings.
--- the companies populating a Stock market, along with their connections, can be effectively modeled through a directed network, where the nodes represent the companies, and the links indicate the ownership. This paper deals with this theme and discusses the concentration of a market. A cross-shareholding matrix is co…
Paper fine-tunes a language model to predict long-term stock buy signals.
problem Predicting long-term stock price movements with narrative text.
method Fine-tuning a small language model on 10-K reports for buy/sell decisions.
result Buy signals generated from 10-K text are most precise at 6 and 9 months, providing 4.8-9% improvement over random selection.
In this paper we examine inefficiencies and information disparity in the Japanese stock market. By carefully analysing information publicly available on the internet, an `outsider' to conventional statistical arbitrage strategies--which are based on market microstructure, company releases, or analyst reports--can never…
This project demonstrated a methodology to estimating cooperate credibility with a Natural Language Processing approach. As cooperate transparency impacts both the credibility and possible future earnings of the firm, it is an important factor to be considered by banks and investors on risk assessments of listed firms.…
Predict stock price movements using financial data and news articles with LLMs.
problem Predicting stock price movements using financial data and news articles.
method Combining financial data and news articles, employing pre-trained LLMs, and using retrieval augmentation techniques.
result Predicted stock price movements with a weighted F1-score of 58.5% and 59.1%.
New attacks inflate earnings while reducing fraud scores, potentially millions at stake.
problem Manipulating financial reports to hide distress and gain.
method Maximum Violated Multi-Objective (MVMO) attacks that adapt search direction.
result Inflation of earnings by 100-200% while reducing fraud scores by 15% in 50% of cases.
New algorithm improves insurance company's asset allocation decisions.
problem Strategic asset allocation with multiple objectives (risk, return, solvency, distance to current portfolio).
method Exact multi-objective optimization algorithm incorporating four objectives.
result Significant improvement in portfolio quality and decision-making process.
Business taxonomies are indispensable tools for investors to do equity research and make professional decisions. However, to identify the structure of industry sectors in an emerging market is challenging for two reasons. First, existing taxonomies are designed for mature markets, which may not be the appropriate class…
This paper analyzes financial sentiment using LLMs and FinBERT, improving accuracy with few-shot examples.
problem Financial sentiment analysis for market evaluation.
method Application of large language models and FinBERT, with focus on prompt engineering and few-shot learning.
result GPT-4o achieves similar sentiment classification accuracy to FinBERT with fewer examples.
Study evaluates digital transformation impact on financial performance using LLMs.
problem Measuring and understanding the impact of digital transformation on financial performance.
method Constructed DT indicators from company reports; analyzed effects of different digital technologies.
result Digital transformation improves financial performance, but varies by technology.
Analyzes systemic risk in European insurance sector using MST and deltaCoVaR.
problem Assessing systemic risk in European insurance sector over 2005-2019.
method Minimum spanning trees (MST) and deltaCoVaR measure.
result The contribution to systemic risk varies by company's centrality in MST.
The study analyzes how bonus-malus systems and delayed claims settlement affect insurance companies' financial stability.
problem Analyzing the impact of bonus-malus systems and delayed claims settlement on insurance companies' financial stability.
method Examined a discrete-time risk model with time-varying premiums, evaluating two types of claims and settlement delays.
result Delayed settlement of by-claims leads to lower ruin probabilities under specific assumptions.
Textual data predicts electricity consumption and weather.
problem Lack of textual data in time series prediction models.
method Used TF-IDF and neural word embeddings to predict time series from text.
result Textual data can predict time series with sufficient accuracy.
In this paper we study data from the yearly reports the four major Swedish non-life insurers have sent to the Swedish Financial Supervisory Authority (FSA). We aim at finding marginal distributions of, and dependence between, losses on the five largest lines of business (LoBs) in order to create models for Solvency Cap…
Study assesses environmental management accounting practices in Bangladesh.
problem Low environmental management accounting practices in Bangladeshi manufacturing companies.
method Developed a compliance checklist and evaluated practices using binary scoring.
result Environmental management accounting practices are poor in Bangladeshi manufacturing companies.
We report the proof that the expression of extended Gibrat's law is unique and the probability distribution function (pdf) is also uniquely derived from the law of detailed balance and the extended Gibrat's law. In the proof, two approximations are employed that the pdf of growth rate is described as tent-shaped expone…
We report the first, to the best of our knowledge, hand-in-hand collaboration between human rights activists and machine learners, leveraging crowd-sourcing to study online abuse against women on Twitter. On a technical front, we carefully curate an unbiased yet low-variance dataset of labeled tweets, analyze it to acc…