Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

35810 · Jun 202619922001200920172026
48 results for disclosures

Model shows disclosure reduces trading costs in oligopolistic markets.

problem Reducing trading costs in oligopolistic markets with imperfect competition.
method Developed a multi-period Kyle-type model with mandatory disclosure and imperfect competition, proving existence and uniqueness of a linear equilibrium.
result Disclosure lowers trading costs by reducing price impact, and its marginal benefit is larger when competition is weak.

Research shows higher damages may encourage more disclosure in corporate disputes.

problem How to resolve disputes over undisclosed material events in a way that encourages voluntary disclosure.
method Dynamic continuous-time model of management's equilibrium disclosure decision.
result Increased damages may lead to an endogenous increase in voluntary disclosure.

New distress dictionary improves bankruptcy prediction from disclosure text.

problem Bankruptcy prediction from financial disclosures.
method Proposes a distress dictionary based on managers' sentences, quantifies linguistic features, and builds predictive models.
result Predictive models based on the distress dictionary outperform existing methods.

Model analyzes how firms balance full disclosure with selective disclosure to maintain a good reputation.

problem Managing reputation in financial markets through voluntary disclosure.
method Developed a dynamic model with two disclosure strategies: candid and sparing, using a piecewise-deterministic model.
result Firms are rewarded for full disclosure but may switch to selective disclosure to avoid potential downgrades.

Study finds companies react negatively to material cybersecurity incident disclosures.

problem Understanding market reactions to cybersecurity incidents.
method Examined daily stock price movements of companies disclosing material cybersecurity incidents.
result Companies tend to experience negative price reactions after disclosing material cybersecurity incidents.

Aggregates diverse zero-shot LLM outputs for better corporate disclosure classification.

problem Combining varied zero-shot LLM predictions for improved stock return prediction.
method Multi-prompt framework with three fixed zero-shot LLM classifiers, logistic meta-classifier aggregation.
result Aggregated model outperforms single classifiers and baseline models, increasing balanced accuracy from 0.566 to 0.606.

Study uses LLM to extract and compare segment disclosures from financial filings.

problem Challenges in completeness and comparability of segment disclosures in financial reports.
method Developed a large language model framework to extract and preserve segment information from Form 10-K filings.
result The LLM accurately extracts segment-level information and addresses cross-period knowledge questions.

Study examines value relevance of oil and gas reserve disclosures in London Stock Exchange.

problem Uncertainty in oil and gas reserves poses accounting challenges for investors.
method Empirical analysis using archival data and multifactor framework.
result Changes in reserves and their components are associated with share returns, but insignificantly due to oil price and longitudinal effects. Quality of disclosures positively impacts share returns.

Study shows cognitive load impacts financial market efficiency, especially for less sophisticated investors.

problem Cognitive load's effect on financial market information processing.
method Developed a theoretical framework and tested it with exogenous disclosure complexity variation.
result Cognitive load significantly impairs price discovery, particularly for less sophisticated investors.

ChatGPT can summarize corporate disclosures more concisely and effectively, improving stock market reactions.

problem Information asymmetry and inefficiency in stock markets due to bloated disclosures.
method Comparing ChatGPT-generated summaries to original disclosures, analyzing their impact on stock market reactions.
result ChatGPT-generated summaries are more effective at explaining stock market reactions to disclosed information.

In this paper, we present a multi-period trading model in the style of Kyle (1985)'s inside trading model, by assuming that there are at least two insiders in the market with long-lived private information, under the requirement that each insider publicly discloses his stock trades after the fact. Based on this model, …

2011-03-04abs ↗pdf ↗

Evidence acquisition costs influence disclosure behavior and preference.

problem How evidence acquisition costs affect disclosure behavior and preference.
method Analyzes sender-receiver interactions with covert and overt evidence acquisition, varying certification costs.
result Equilibria converge to the Pareto-worst free-learning equilibrium as costs vanish, and receivers prefer covert to overt acquisition.

CAI automates extraction and validation of corporate GHG emission metrics.

problem Manual extraction of corporate GHG emission metrics is labor-intensive and error-prone.
method CAI uses LLMs to automate extraction and validation of metrics from corporate disclosures.
result CAI improves data collection efficiency and accuracy by automating the process.

A new metric GNQ audits LLMs for privacy risks during training.

problem Auditing LLMs for privacy risks during training is computationally hard.
method Gradient Uniqueness (GNQ) metric derived from gradient descent, BS-Ghost GNQ for efficiency.
result GNQ successfully predicts sequence extractability and reveals risk heterogeneity.

In this paper, we present a multi-period trading model by assuming that traders face not only asymmetric information but also heterogenous prior beliefs, under the requirement that the insider publicly disclose his stock trades after the fact. We show that there is an equilibrium in which the irrational insider camoufl…

2011-05-12abs ↗pdf ↗

Paper creates transparent, safe synthetic data from coarsened margins.

problem Creating synthetic data that maintains original relationships and is safe from disclosure.
method Defining and curating margins, applying SDC, coarsening counts, and using IPF algorithm.
result Synthetic data derived from safe, coarsened margins maintains original relationships.

FinAI-BERT classifies AI disclosures in financial reports with high accuracy.

problem Systematic detection of AI-related disclosures in financial reports.
method Fine-tuned transformer-based model on a curated dataset.
result Achieved near-perfect classification performance (99.37% accuracy).

The study finds that firm membership in flagship indices and TCFD endorsement are strong predictors of a wider Disclosure-Performance Gap.

problem The Aggregate Confusion hypothesis and the measurement of greenwashing in environmental disclosures.
method The study uses a Disclosure-Performance Gap (DPG) model to measure the divergence between voluntary environmental disclosures and realised emissions performance for 200 large European firms. The model selection process involved multiple stages and robust standard errors.
result Firm membership in flagship indices and TCFD endorsement are strong predictors of a wider gap, while renewable energy use and environmental capital expenditure significantly narrow the gap.

Study identifies a Strategic Gap in market efficiency due to AI-driven timing and complexity in disclosure.

problem Market inefficiency due to structural influence of disclosure timing and complexity.
method Introduces Autonomous Disclosure Regulator, a multi-node AI framework to audit disclosure complexity and unpredictability.
result Companies use confusing language and unpredictable timing to slow down market learning, creating a 60% Structural Gap.

Weak predictability of stock price movement 2 days after annual report disclosure.

problem Predicting stock price movement after annual report disclosure.
method Used various models including decision tree, logistic regression, random forest, neural network, prototypical networks; used financial indicators from EastMoney.
result Maximum accuracy and precision of stock price movement prediction is around 59.6% and 0.56 respectively, with random forest performing best.

The study improves sentiment analysis of 10-K filings, revealing aggregation effects on accuracy and correlation with market outcomes.

problem Lack of sentiment analysis for 10-K filings, particularly for risk disclosures.
method Supervised lexicon-learning approach applied to 10-K filings and Item 1A risk-factor sections, trained against return and volatility labels at different levels of aggregation.
result Sentiment analysis of Item 1A sections performs better at the individual-firm level, while full-filing text is more accurate at sector and portfolio levels.

Fund2Persona creates personalized financial advisor personas from fund data, improving investment advice.

problem Lack of consistent advisor expertise and difficulty in encoding it in LLM systems.
method Grounds financial advisor personas in fund disclosures, market context, and manager commentary through an agentic actor--scorer--patcher loop.
result Personas better recover portfolio decisions and manager interpretation than generic baselines.

Analyzes how uncertainty in financial networks affects stability.

problem Understanding how uncertainty in financial networks impacts stability.
method Introduced a minimal stochastic dynamical model of the interbank network with linear interactions. Derived the interaction correction to the stress expectation and studied it on the short-medium timescale.
result Interactions increase the stress expectation on average, highlighting the importance of disclosure.

Fund2Persona creates personalized financial advisor personas from fund data, improving investment advice and manager interpretation.

problem Lack of consistent and specific financial advisor expertise in personalized investment advice.
method Grounds financial advisor personas in fund disclosures, holdings transitions, market context, and manager commentary through an agentic actor--scorer--patcher loop.
result Personas better recover portfolio decisions and grounded manager interpretation than generic baselines.

Developed a new method to generate synthetic data while protecting privacy.

problem Protecting privacy of human participant data while making it publicly accessible.
method A multi-step framework based on Classification and Regression Trees and an original distance-based filtering.
result Satisfactory protection against attribute disclosure attacks and formal prevention of membership disclosure attacks.

We propose a categorical data synthesizer with a quantifiable disclosure risk. Our algorithm, named Perturbed Gibbs Sampler, can handle high-dimensional categorical data that are often intractable to represent as contingency tables. The algorithm extends a multiple imputation strategy for fully synthetic data by utiliz…

2013-12-18abs ↗pdf ↗

Synthetic tabular data synthesis models balance utility and risk.

problem Generating synthetic tabular data for regulated domains.
method Latent flow models with various learning targets, paths, and sampling methods.
result Velocity and posterior matching objectives yield higher utility, while score and noise matching achieve lower risk.

Narrative disclosures in 10-K filings improve bankruptcy prediction beyond accounting ratios.

problem Traditional bankruptcy prediction models rely on accounting ratios, which may not capture early warning signals.
method Developed a PB Stress Score based on distress-specific language in 10-K narratives, evaluated against accounting and dictionary benchmarks.
result Adding the PB Stress Score increases AUC from 0.8323 to 0.9019 and improves top-decile bankruptcy capture from 44.12% to 64.71%.

The study finds that supply chain information from LLM embeddings improves stock returns predictions.

problem Predicting stock returns using textual information from annual reports.
method Combining LLM embeddings of annual reports with supply chain knowledge graph propagation.
result Network-augmented embeddings significantly predict stock returns with a Sharpe ratio of 0.86 and alpha of 7.27%.

LLMs help less-resourced researchers access costly data.

problem Unequal access to costly datasets limits research contributions.
method RAG framework with GPT-4o-mini for automated data collection.
result LLMs can collect CEO pay ratios and CAMs from corporate disclosures with high accuracy and low cost.

In this paper, we propose FedGP, a framework for privacy-preserving data release in the federated learning setting. We use generative adversarial networks, generator components of which are trained by FedAvg algorithm, to draw privacy-preserving artificial data samples and empirically assess the risk of information dis…

2019-10-18abs ↗pdf ↗