Paper introduces a trading agent using LLMs for risk assessment and trading recommendations.
problem Developing a trading agent that can handle financial risks effectively.
method Extending CPPO algorithm with LLM-generated risk assessment and trading signals from financial news.
result Backtesting shows improved performance of the trading agent compared to benchmarks.
LLMs can help explain credit risk models but not autonomously.
problem Leveraging LLMs for post-hoc explainability in credit risk models.
method Comparison of LLM outputs with SHAP and coefficient-based attributions on three LMs.
result LLMs reliably preserve feature-importance rankings but poorly align with autonomous explanations.
RiskLabs uses LLMs to predict financial risks from multimodal data.
problem Financial risk prediction using AI techniques.
method Integrates multimodal financial data (textual, vocal, time series, news) into LLMs for prediction.
result Empirical results show effectiveness in forecasting market volatility and variance.
New benchmarks focus on LLM risk in finance, not just accuracy.
problem Standard benchmarks ignore LLM financial risks, leading to unsafe deployment.
method Three-level agenda: model, workflow, and system stress-testing.
result Hidden weaknesses in LLMs are missed by traditional benchmarks.
The study analyzes how large language models form and express investor risk profiles.
problem Understanding how large language models (LLMs) form and express investor risk profiles.
method Examined three LLMs (GPT, Gemini, and Llama) and assessed their responses to a standardized risk questionnaire under varying prompts.
result LLMs generally form long-term investment profiles, but they exhibit different risk tolerance levels.
Agentic LLMs improve trading by estimating market risk.
problem Lack of principled model-building step in agentic frameworks for finance.
method Developed an agentic system using LLMs to discover stochastic differential equations for financial time series.
result Model-informed trading strategies outperform standard LLM-based agents, improving Sharpe ratios.
This paper proposes a new portfolio allocation method using LLMs to outperform traditional strategies.
problem Persistent tradeoff between risk and return in portfolio management.
method Follow-the-leader approach with sentiment-based trade filtering and LLM-driven hedging.
result Empirical results show a 69% increase in annualized returns and 119% in Sharpe ratio compared to SPY buy-and-hold.
LLM generates coherent macroeconomic stress scenarios for portfolio risk assessment.
problem Macro-financial stress testing and portfolio risk assessment using traditional methods.
method Hybrid prompt-RAG pipeline combining structured prompting and retrieval of country fundamentals and news.
result LLM-generated scenarios yield stable tail-risk amplification with limited sensitivity to retrieval choices.
New framework assesses LLM security risks in BFSI.
problem Lack of domain-specific security evaluation for LLMs in BFSI.
method Risk-aware evaluation framework combining taxonomy, automated red-teaming, and ensemble judging.
result Higher decoding stochasticity and adaptive interaction lead to more severe disclosures.
This paper reviews LLMs for credit risk assessment, creating a taxonomy.
problem Assessing credit risk using financial text analysis.
method Systematic review of 60 papers, focusing on model architectures, data types, and explainability mechanisms.
result Developed a taxonomy of LLM-based credit risk models.
LLM trading agents show risk feedback can improve alignment without fine-tuning.
problem Aligning LLM trading agents with financial risk.
method TradeArena testbed, risk reports, execution simulation, memory replay.
result Risk feedback can improve alignment without fine-tuning, but not universally.
Proposes a method to choose thresholds for LLM evaluation metrics.
problem Ensuring reliable large language models (LLMs) with correct threshold selection.
method Identify risks, stakeholders' risk tolerance, and use ground-truth data to determine thresholds.
result Demonstrates a concrete example with the Faithfulness metric and HaluBench dataset.
Method estimates LLM error rates using Pareto optimization.
problem Quantifying error rates in text-generating models.
method Pareto optimization for generating risk scores.
result Risk scores correlate well with true error rates.
Simple online monitor detects unsafe LLM outputs.
problem LLMs generate unsafe outputs despite training.
method Thresholding external verifier signal to decide alarms.
result Simple design competitive with advanced methods.
LLMs can fail to maximize aligned values even after training, due to irrational reasoning.
problem Value misalignment in LLMs' reasoning despite training alignment.
method Formalized rational value risk and decomposed estimation error.
result Rational value risk is widespread and cannot be fully eliminated.
Develops a framework for quantifying agentic AI model risk using LLM-inferred Bayesian state filters.
problem Quantifying the risk of agentic AI systems due to uncertain beliefs and actions.
method Representing the system as a partially observed Markov decision process with latent states, Bayesian belief updates, control-dependent losses, and tail-risk functionals.
result Develops a rigorous framework for separating uncertainty quantification from risk measurement.
AlphaSharpe uses LLMs to improve financial metrics robustness and predictive power.
problem Traditional financial metrics struggle with robustness and generalization in volatile markets.
method Iterative optimization of financial metrics using LLMs, including crossover, mutation, and evaluation.
result AlphaSharpe discovers enhanced risk-return metrics with 3x predictive power and 2x portfolio performance.
LLMs struggle to outperform markets over long periods and diverse stocks.
problem Overstated effectiveness of LLM-based investing strategies due to biases.
method FINSABER framework for systematic backtests over two decades and 100+ symbols.
result Previously reported LLM advantages deteriorate significantly under broader evaluation.
Framework ensures alignment between humans and machines in LLMs.
problem Human-machine misalignment in LLMs scoring mechanisms.
method Lightweight calibration framework for blackbox models.
result Provably guarantees alignment between humans and machines.
LLM sandbox and persona dynamics create unethical reality gaps that shift risk to users.
problem Ethical issues arise from LLMs generating reality gaps that shift risk to uninformed users.
method Analyzes the ethical implications of LLM sandbox and persona dynamics, comparing them to financial regulation and compliance.
result Active generation of reality gaps is unethical as it shifts epistemic risk to users.
Hybrid model uses LLM to build transparent Bayesian networks for trading decisions.
problem Rigorous and transparent reasoning required in financial trading, especially for options strategies.
method Combines LLM strengths with Bayesian Networks, using LLM to construct context-specific networks and select relevant data.
result Empirically, the hybrid system outperforms market benchmarks with superior risk-adjusted performance.
A new metric GNQ audits LLMs for privacy risks during training.
problem Auditing LLMs for privacy risks during training is computationally hard.
method Gradient Uniqueness (GNQ) metric derived from gradient descent, BS-Ghost GNQ for efficiency.
result GNQ successfully predicts sequence extractability and reveals risk heterogeneity.
A system for supervising decentralized finance risks using LLMs and structured evidence.
problem Supervising decentralized finance risks
method Forecast-grounded agentic supervision system
result Developed a system that scores tickets against a regulator-aligned ground truth and false-intervention rate.
Computer science scans LLMs to understand and manipulate their economic forecasts.
problem Understanding and controlling the reasoning of large language models in economics.
method Brain scanning techniques applied to LLMs to identify and manipulate underlying concepts.
result LLMs can be steered to generate forecasts with specific biases, allowing for correction or simulation.
Trading-R1 uses LLMs for financial trading, improving risk-adjusted returns.
problem Lack of interpretability and trust in AI for finance.
method Supervised fine-tuning and reinforcement learning with a curriculum.
result Improved risk-adjusted returns and lower drawdowns compared to other models.
Framework uses LLMs to automate strategy finding in quantitative finance.
problem Brittleness of traditional deep learning models in financial applications.
method Three-stage framework with prompt-engineered LLMs, multimodal agent-based evaluation, and dynamic weight optimization.
result Robust performance in Chinese & US markets, superior risk-adjusted performance.
DeXposure-Claw supervises decentralized finance risks by grounding LLM decisions in evidence.
problem Weak evidence leads to over-interventions by general-purpose LLM agents in decentralized finance.
method DeXposure-Claw uses a graph time-series foundation model to forecast exposure networks, turning forecasts into alerts and constraining escalation with data-health gates.
result DeXposure-Claw reduces false alarms and improves regulator alignment in decentralized finance risk supervision.
CSA fills a gap in RLVR-trained LLM deployment by providing anytime-valid selective risk control.
problem Deployment of RLVR-trained LLMs in regulated organizations requires a safety certificate for every round without waiting for long-run averages.
method CSA uses a (test statistic, validity guarantee, deployment rule) framework to fill the gap, maintaining a Ville-type e-process per threshold on a Bonferroni grid.
result CSA provides the first anytime-valid selective risk control for RLVR-trained LLMs, matching the long-run average certification rate and satisfying pathwise validity and non-refusing deployment on every cell.
ChatGPT improves momentum strategies by analyzing news data.
problem Improving risk-adjusted returns in systematic investing.
method Combining LLMs with daily equity returns and news data to predict stock momentum.
result LLM-enhanced momentum strategies outperform benchmarks in Sharpe and Sortino ratios.
CROQ optimizes LLM decision-making by narrowing down choices and improving accuracy.
problem Uncertainty in LLM outputs poses risks in high-stakes domains.
method Conformal prediction (CP) and optimization (CP-OPT) to minimize prediction set sizes.
result CROQ improves LLM accuracy, especially with CP-OPT.
LLMs overestimate stock returns and are less accurate at predicting extreme outcomes.
problem Behavioral biases in LLMs' stock return forecasts.
method Comparison of LLM forecasts with crowd-sourced estimates and historical data.
result LLMs overestimate stock returns and are less accurate at predicting extreme outcomes.
FinHEAR combines LLMs with human expertise for better financial decision-making.
problem Challenges in financial decision-making for language models.
method Multi-agent framework with specialized LLMs for historical analysis, event interpretation, and expert retrieval.
result FinHEAR outperforms baselines in financial tasks with higher accuracy and risk-adjusted returns.
Paper proposes FinAR-Bench to evaluate LLMs in financial analysis tasks.
problem Inaccurate financial analysis by LLMs leading to investment and regulatory issues.
method Proposes FinAR-Bench, a benchmark dataset with three steps: key info extraction, financial indicator calculation, and logical reasoning.
result LLMs perform better in key info extraction and indicator calculation but struggle with logical reasoning.
FinPT uses large pretrained models to predict financial risks.
problem Outdated algorithms and lack of open financial benchmarks.
method Profile Tuning on large pretrained foundation models.
result Demonstrated effectiveness on FinBench datasets.
TRIBE model uses LLMs to simulate human trading behavior in bond markets.
problem Complexities in decentralized bond market transactions.
method Agent-based model augmented with LLMs to simulate human-like decision-making.
result Slight trade aversion in LLMs can lead to complete market collapse.
LLMs mimic human traders in finance, but not as much as expected.
problem Evaluating how LLMs behave in financial markets.
method Adapted experimental design with LLMs and human traders, analyzed in single and mixed model settings.
result LLMs tend to price assets near their fundamental value, but not as much as humans, and show less trading strategy variance.
Fine-tuning LLMs on privacy-sensitive data introduces privacy risk, and synthetic data audits can quantify this risk.
problem Fine-tuning LLMs on privacy-sensitive data introduces privacy risk.
method Generate synthetic canaries via high-temperature sampling from LLMs.
result Synthetic canaries are high-influence outliers that ensure strong audits.
New method certifies risks of LLM outputs, improving accuracy and reliability.
problem Uncertain and incorrect outputs from large language models.
method Information-lift certificates using PAC-Bayes bounds and skeleton design.
result Achieves 77.0% coverage at 2% risk, outperforming baselines.
Paper analyzes systematic jump risk around the clock using news narratives.
problem Identifying and managing priced risks in real-time market conditions.
method Combining high-frequency market data with news narratives classified by an LLM.
result Significant heterogeneity in risk premia, with macroeconomic news commanding the largest premium.
This paper combines LLMs with RL for better trading strategies.
problem Myopic behavior and opaque policies in RL for trading.
method LLMs generate strategic trading advice to guide RL agents.
result LLM-guided RL agents outperform unguided RL in return and risk metrics.
Paper uses LLMs for sector allocation, showing better returns.
problem Automated trading sector allocation inefficiencies.
method Systematic analysis of macroeconomic data and sentiment.
result LLM-based sector allocation outperforms traditional strategies.
TradingAgents uses LLM-powered multi-agent framework for financial trading.
problem Lack of collaborative dynamics in multi-agent financial trading systems.
method Inspired by real-world trading firms, TradingAgents features specialized LLM-powered agents and a risk management team.
result Framework outperforms baseline models in trading performance metrics.
LLMs improve financial analysis by processing large data sets.
problem Traditional financial analysis methods struggle with large data volumes.
method Integrating LLMs for enhanced data processing and analysis.
result LLMs offer new capabilities for real-time financial decision-making.
Study uses LLMs to simplify financial regulation interpretation.
problem Complex financial regulations are hard to interpret and implement.
method Developed prompts to guide LLMs in extracting key information from regulations.
result GPT-4 outperforms other LLMs in processing and executing regulatory requirements.
Valid certifies LLMs' domain adherence, bounding out-of-domain behavior.
problem Adversarial susceptibility of LLMs to generate out-of-domain outputs.
method VALID approach providing adversarial bounds as a certificate.
result Validates LLMs' domain adherence with meaningful certificates.
LiveTradeBench evaluates LLMs in live trading environments.
problem Static benchmarks fail to assess real-world trading ability.
method Live data streaming, portfolio management abstraction, multi-market evaluation.
result LLMs show distinct portfolio styles and adapt to live signals.
Hybrid method uses LLM to filter lead-lag relationships in prediction markets.
problem Challenges in discovering robust lead-lag relationships in prediction markets due to spurious correlations.
method Two-stage approach: statistical Granger causality followed by LLM semantic re-ranking.
result LLM-based method outperforms statistical baseline, increasing win rate and reducing average loss magnitude.
This paper uses CausalGANs and RL with LLM to predict bond yields.
problem Challenges in financial bond yield forecasting due to data scarcity and market conditions.
method Proposes a novel framework combining CausalGANs, RL, and LLM for synthetic data generation and trading signals.
result Improves forecasting performance over existing methods with low Mean Absolute Error.