New benchmarks focus on LLM risk in finance, not just accuracy.
problem Standard benchmarks ignore LLM financial risks, leading to unsafe deployment.
method Three-level agenda: model, workflow, and system stress-testing.
result Hidden weaknesses in LLMs are missed by traditional benchmarks.
Optimal benchmark design varies based on costs in financial manipulation.
problem Manipulation of price benchmarks in finance.
method Analyzes empirical pattern and cost structures to determine optimal benchmark design.
result The optimal benchmark depends on the relative sizes of fixed and variable costs.
Survey of financial LLMs in finance.
problem Limited research on Financial LLMs (FinLLMs).
method Chronological overview of PLMs, comparison of techniques, performance evaluations, and advanced tasks.
result Compilation of accessible datasets and benchmarks for AI research in finance.
Study benchmarks classical models over quantum in DeFi yield prediction.
problem Accurate yield and performance forecasting for DeFi liquidity allocation.
method Benchmarked six models on Curve Finance pools' historical data.
result Classical models, especially XGBoost, outperform quantum models.
IPO Finance Agent extends Finance Agent v2 for SpaceX S-1 filings, improving accuracy and cost-efficiency.
problem Evaluating IPO due diligence tasks with long-form documents.
method Extended task domain, improved agentic harness with contextual retrieval, automated rubric generation.
result Best-performing model reaches 79.8% accuracy, cost-efficient model at 77.2% with 0.05 USD per query.
Study proposes new methods to calculate probabilistic benchmarks in noisy data.
problem Identifying opportunities for improvement in comparable units with noisy data.
method 2-step methodology involving undersampling and relevance vector machine.
result Higher discrimination power achieved with macro-economic environment variables.
IPO Finance Agent evaluates LLMs on SpaceX IPO due diligence, surpassing Finance Agent v2.
problem Evaluating language models on financial tasks like IPO due diligence.
method Introducing IPO Finance Agent and an evaluator-optimizer pipeline.
result The best-performing model reaches 79.4% accuracy at 0.30 USD per query.
FinTMMBench benchmarks RAG systems for finance tasks across multiple data types and time periods.
problem Evaluating temporal-aware multi-modal retrieval augmented generation in finance.
method TMMHybridRAG method that converts and integrates data from various modalities and temporal information.
result Demonstrated effectiveness of TMMHybridRAG in diverse financial analysis tasks.
NMIXX fine-tunes embeddings for finance, outperforming general models in Korean.
problem Financial embeddings struggle in low-resource languages like Korean.
method Fine-tuned with 18.8K triplets, hard negatives, and translations.
result NMIXX achieves gains of +0.10 on English FinSTS and +0.22 on KorFinSTS.
We discuss suitable classes of diffusion processes, for which functionals relevant to finance can be computed via Monte Carlo methods. In particular, we construct exact simulation schemes for processes from this class. However, should the finance problem under consideration require e.g. continuous monitoring of the pro…
FinSurvival provides a large-scale financial survival modeling benchmark.
problem Lack of large-scale, realistic, and freely available datasets for benchmarking AI survival models.
method Derived 16 survival modeling tasks from cryptocurrency lending data using an automated pipeline.
result Demonstrated that existing AI survival models are not well-suited for these challenging tasks.
Framework uses LLMs to automate strategy finding in quantitative finance.
problem Brittleness of traditional deep learning models in financial applications.
method Three-stage framework with prompt-engineered LLMs, multimodal agent-based evaluation, and dynamic weight optimization.
result Robust performance in Chinese & US markets, superior risk-adjusted performance.
LLM Pro Finance Suite enhances financial NLP with instruction-tuned models.
problem Limited NLP capabilities for financial tasks in generalist models.
method Instruction-tuned large language models fine-tuned on financial data.
result Consistent improvement over state-of-the-art baselines in finance tasks.
Study builds a Japanese financial-specific LLM through continual pre-training.
problem Lack of domain-specific Japanese financial LLMs.
method Continual pre-training on Japanese financial-focused datasets using a base Japanese LLM.
result Tuned model outperforms original model on Japanese financial benchmarks.
FinRL-Meta offers market environments and benchmarks for financial reinforcement learning.
problem Challenges in creating high-quality market environments and benchmarks for financial reinforcement learning.
method DataOps paradigm, automatic pipeline, community-wise competitions, Jupyter/Python demos.
result Openly accessible FinRL-Meta library for data-driven financial reinforcement learning.
We develop a statistical framework to benchmark and select large language models based on their risks.
problem Benchmarking and selecting large language models based on their associated risks.
method A distributional framework using first and second order stochastic dominance, linked to mean-risk models in finance.
result Formalizes a risk-aware approach for model selection, balancing risk and utility.
RL applied to finance tasks, highlighting challenges and future directions.
problem Decision-making tasks in finance using RL.
method Meta-analysis of RL applications, identifying challenges and proposing future directions.
result Challenges in RL performance and future research directions.
Paper proposes FinAR-Bench to evaluate LLMs in financial analysis tasks.
problem Inaccurate financial analysis by LLMs leading to investment and regulatory issues.
method Proposes FinAR-Bench, a benchmark dataset with three steps: key info extraction, financial indicator calculation, and logical reasoning.
result LLMs perform better in key info extraction and indicator calculation but struggle with logical reasoning.
Paper introduces NumLLM for better financial text understanding with numeric variables.
problem Poor performance of existing financial large language models in numeric financial text.
method Constructed financial corpus, fine-tuned with LoRA modules, merged into foundation model.
result NumLLM achieves best performance on financial question-answering benchmark, especially with numeric questions.
FINCH dataset enables financial Text-to-SQL tasks, improving model evaluation.
problem Lack of a large-scale financial dataset for Text-to-SQL tasks.
method Curated financial dataset (FINCH), benchmarking reasoning and language models.
result Proposed FINCH Score for more accurate financial model evaluation.
REALFIN benchmarks financial reasoning by removing implicit assumptions, revealing model weaknesses.
problem Models struggle when implicit assumptions are missing, leading to incorrect answers.
method Developed a bilingual benchmark that systematically removes essential premises from financial questions.
result General-purpose models over-commit, while finance-specialized models fail to identify missing premises.
Study interprets deep learning models for Heston model in finance.
problem Interpreting black-box deep learning models in finance.
method Investigated Heston model using local and global strategies from cooperative game theory.
result Shapley values can effectively explain neural networks and improve model architecture selection.
SusGen-GPT improves financial NLP and ESG report generation.
problem Lack of advanced NLP tools for finance and ESG domains.
method Developed SusGen-30K dataset and SusGen-GPT models.
result Achieved state-of-the-art performance in financial NLP tasks.
Paper uses DDQN for trading assets, showing better performance than market benchmarks.
problem Improving financial trading strategies using AI.
method Double Deep Q-Network (DDQN) algorithm for trading multiple assets.
result Trading agent outperformed market benchmarks and achieved higher net asset value.
Paper studies pricing and hedging of nonreplicable insurance contracts using benchmark-neutral approach.
problem Pricing and hedging of long-term insurance contracts like variable annuities.
method Benchmark-neutral pricing framework using stock growth optimal portfolio as numéraire.
result Prices can be significantly lower than risk-neutral ones, offering attractive long-term risk-management.
Deep learning models struggle with new data in stock price trend prediction.
problem Stock price trend prediction using Deep Learning models.
method Examination of fifteen state-of-the-art DL models on LOB data, using LOBCAST framework.
result All models show significant performance drop with new data, questioning their market applicability.
Study benchmarks LLMs in portfolio optimization tasks.
problem Evaluate financial decision-making of LLMs.
method Mathematically explicit portfolio optimization problems with multiple-choice questions.
result Distinct performance patterns among LLMs in different financial tasks.
LOB-Bench benchmarks generative AI for financial data, outperforming traditional models.
problem Lack of consensus on evaluating generative AI models for financial data.
method Python-based benchmark with LOB statistics and market impact metrics.
result Generative autoregressive models outperform traditional models in LOB data.
Study examines stock price reactions to Texas winter storm power outages.
problem Impact of natural disasters on stock market values.
method Used four benchmark models to measure abnormal returns.
result Firms experienced significant stock price drops after the Texas winter storm.
The paper tackles fVaR prediction methods in finance.
problem Predicting future values at risk (fVaR) in finance.
method Various methods including Nested MC-empirical quantile, percentiles from distributions, quantile regressions, and limited inner simulations.
result Improved methods for predicting fVaRs, including those that are computationally efficient.
Quantum computer optimizes investment portfolios, outperforming traditional methods.
problem Minimizing risk while meeting return and budget constraints in investment portfolios.
method Used D-Wave quantum annealer and hybrid solvers to solve Portfolio Optimization problem.
result D-Wave quantum solution performs close to traditional commercial solvers for tested problem sizes.
This study evaluates zero-shot LLMs in finance, finding ChatGPT performs well but fine-tuned models are better.
problem Evaluating zero-shot LLMs in financial tasks.
method Comparison of ChatGPT and fine-tuned models on annotated data.
result Fine-tuned models generally outperform zero-shot LLMs.
Deep RL controller outperforms market making benchmarks in a Hawkes process model.
problem Optimal market making in financial markets.
method Deep reinforcement learning on a Hawkes process-based simulator.
result Deep RL controller outperforms benchmarks in various risk-reward metrics.
CSTS benchmarks time series clustering by evaluating correlation structures.
problem Lack of validated ground truth for objectively assessing clustering quality.
method Synthetic benchmark CSTS for evaluating correlation structures in multivariate time series data.
result CSTS enables precise diagnosis of methodological limitations in correlation-based time series clustering.
This review analyzes RL in finance, highlighting its advantages and challenges.
problem Complex financial decision-making problems where traditional methods fail.
method Systematic review of 167 articles from 2017-2025, focusing on market making, portfolio optimization, and algorithmic trading.
result RL offers advantages over traditional methods, particularly in market making, but challenges remain.
Quantum reservoir computing improves volatility forecasting.
problem Forecasting realized volatility in finance.
method Quantum reservoir computing with Ising Hamiltonian and feature selection.
result Quantum reservoir computing outperforms benchmarks in volatility forecasting.
CTBench benchmarks cryptocurrency time series generation for trading applications.
problem Lack of comprehensive benchmarks for cryptocurrency time series generation.
method Developed a comprehensive benchmark extsf{CTBench} with 13 metrics across 5 dimensions.
result Uncovered trade-offs between statistical fidelity and real-world profitability.
Study proposes optimal risk-aware interest rates for crypto lending protocols.
problem Determining optimal interest rates for decentralized lending protocols to maximize profit and minimize risk.
method Agent-based model, Riccati-type ODEs for linear behaviors, Monte-Carlo estimator and deep learning for nonlinear behaviors.
result Calibrated model shows superior risk-adjusted performance compared to industry-standard interest rate models.
NeuralBeta uses deep learning to estimate beta, outperforming traditional methods.
problem Limitations of traditional beta estimation methods in capturing dynamic beta behavior.
method Neural networks with a new output layer for interpretability.
result NeuralBeta outperforms benchmark methods in dynamic beta estimation.
In this paper we address three main objections of behavioral finance to the theory of rational finance, considered as anomalies the theory of rational finance cannot explain: Predictability of asset returns, The Equity Premium, (The Volatility Puzzle. We offer resolutions of those objections within the rational finance…
In this paper we consider the pricing of variable annuities (VAs) with guaranteed minimum withdrawal benefits. We consider two pricing approaches, the classical risk-neutral approach and the benchmark approach, and we examine the associated static and optimal behaviors of both the investor and insurer. The first model …
Trade finance history traced from medieval origins to modern markets.
problem Evolution and standardization of trade finance products.
method Historical analysis of market structures and regulatory changes.
result Global trade finance market evolved from local to centralized, then decentralized.
Paper reviews the evolution of alpha from human insight to AI-powered systems.
problem Exceeding market benchmarks in finance.
method Five-stage taxonomy integrating representation learning, multimodal data fusion, and LLM agents.
result Unified framework for evaluating and developing next-gen alpha systems.
Given a set of empirical observations, conditional density estimation aims to capture the statistical relationship between a conditional variable x and a dependent variable y by modeling their conditional probability p(y∣x). The paper develops best practices for conditional den…
Auto.gov uses RL to automate DeFi governance, improving security and profitability.
problem Manual DeFi governance is prone to human bias and financial risks.
method Auto.gov employs a deep Q-network reinforcement learning strategy for semi-automated parameter adjustments. result Auto.gov outperforms traditional governance methods by at least 14% in terms of protocol profitability. BloombergGPT is a large language model trained on financial data, outperforming existing models on financial tasks.
problem Lack of specialized large language models for finance.
method Trained on a 363 billion token dataset augmented with 345 billion tokens from general datasets, using a 50 billion parameter model.
result BloombergGPT outperforms existing models on financial tasks without sacrificing performance on general LLM benchmarks.
Decentralized finance uses blockchain for $70B in assets, differing from traditional finance.
problem Ensuring compliance and security in decentralized finance.
method Systematic analysis of legal, economic, security, and privacy aspects.
result Decentralized finance offers unique economic effects and security features.
Alternative finance models from physics for non-equilibrium systems.
problem Inequities of classical finance models in physics-based perspective.
method Physics-based insights for non-equilibrium finance models.
result Alternative models for non-equilibrium finance systems.