Study optimizes compute usage for LLM web agents, improving performance.
problem High compute costs and narrow focus on single-step tasks limit LLM web agents.
method Two-stage pipeline: SFT followed by on-policy RL, with hyperparameter optimization.
result Combining SFT and on-policy RL requires 55% less compute to match peak SFT performance.
FinMem enhances LLM trading agents with layered memory and character design.
problem Developing purpose-driven LLM agents for financial decision-making.
method Integrates layered memory and character design modules into an LLM framework.
result Significantly enhanced trading performance in financial markets.
Enhances anomaly detection in financial markets using AI agents.
problem Manual verification of financial market anomalies is time-consuming and error-prone.
method A multi-agent LLM framework for automated anomaly detection.
result Framework reduces human intervention and improves efficiency and accuracy.
New approach treats coordination as an architectural layer to improve LLM-based multi-agent systems.
problem Coordination defects lead to high failure rates in LLM-based multi-agent systems.
method Treats coordination as a configurable architectural layer separable from agent logic and information access.
result Configurations leave distinguishable signatures, enabling architectural reasoning and Pareto frontiers.
LiveTradeBench evaluates LLMs in live trading environments.
problem Static benchmarks fail to assess real-world trading ability.
method Live data streaming, portfolio management abstraction, multi-market evaluation.
result LLMs show distinct portfolio styles and adapt to live signals.
Survey examines LLMs in financial trading.
problem Using LLMs to outperform professional traders in finance.
method Comprehensive review of current research on LLMs in financial trading.
result LLMs can potentially outperform professional traders in backtesting.
InvestorBench benchmarks LLM agents in financial tasks.
problem Lack of a comprehensive benchmark for LLM-based financial agents.
method Developed a benchmark with diverse financial tasks and datasets.
result Evaluated LLM agents' performance across various financial products and market environments.
MathChat uses LLM agents to solve challenging math problems through conversational problem-solving.
problem Solving math problems expressed in natural language.
method MathChat is a conversational framework combining an LLM agent and a user proxy agent for collaborative problem-solving.
result MathChat improves tool-using prompting methods by 6% on difficult math problems.
CausalGame benchmarks LLM agents' causal thinking in games.
problem Evaluating causal thinking in AI Scientists with LLMs.
method Interactive games with 14 scenarios incorporating selection bias, measurement error, and hidden confounders.
result None of the 30 LLM agents demonstrated reliable causal thinking, with the best model achieving only 68.0% survival.
This paper explores how LLMs can improve pipeline-based conversational agents.
problem Limitations of pipeline-based conversational agents in human-like conversations.
method Investigated LLMs' capabilities in two phases: design and development, and operations.
result LLMs can enhance pipeline-based agents in various tasks like data generation, intent classification, and auto-correction.
TraderTalk uses LLMs to simulate human trading interactions in financial markets.
problem Simulating realistic human trading interactions in financial markets.
method Hybrid ABM with LLM-generated behaviors for detailed conversations.
result Successfully replicates trade-to-order volume ratios in financial markets.
Agentic LLMs improve trading by estimating market risk.
problem Lack of principled model-building step in agentic frameworks for finance.
method Developed an agentic system using LLMs to discover stochastic differential equations for financial time series.
result Model-informed trading strategies outperform standard LLM-based agents, improving Sharpe ratios.
Paper introduces a trading agent using LLMs for risk assessment and trading recommendations.
problem Developing a trading agent that can handle financial risks effectively.
method Extending CPPO algorithm with LLM-generated risk assessment and trading signals from financial news.
result Backtesting shows improved performance of the trading agent compared to benchmarks.
LLMs mimic human traders in finance, but not as much as expected.
problem Evaluating how LLMs behave in financial markets.
method Adapted experimental design with LLMs and human traders, analyzed in single and mixed model settings.
result LLMs tend to price assets near their fundamental value, but not as much as humans, and show less trading strategy variance.
Agent Trading Arena trains LLMs in real-time financial markets to improve numerical reasoning.
problem Limited real-world training for LLMs in financial markets.
method Virtual zero-sum stock market with competitive multi-agent trading.
result LLMs perform better with chart-based visualizations and a reflection module.
Benchmark evaluates LLM trading agents by masking identifiers to prevent memory leaks.
problem Evaluate LLM trading agents without relying on market memory or noise.
method Data-side masking protocol, Barra-style performance attribution framework.
result LLM agents' returns are largely explained by market and style exposure, not stock selection.
TradingAgents uses LLM-powered multi-agent framework for financial trading.
problem Lack of collaborative dynamics in multi-agent financial trading systems.
method Inspired by real-world trading firms, TradingAgents features specialized LLM-powered agents and a risk management team.
result Framework outperforms baseline models in trading performance metrics.
WebGUM learns web navigation from multimodal data, outperforming previous methods.
problem Limited generalization from domain-specific models in web navigation.
method Instruction-following multimodal agent trained on vision-language foundation models.
result Significant improvement in web navigation performance on benchmarks.
ATLAS uses LLMs to adaptively trade by optimizing prompts and coordinating agents.
problem Adapting LLMs for real-time financial decision-making in noisy markets.
method ATLAS integrates structured market data, uses Adaptive-OPRO for prompt optimization, and employs multi-agent coordination.
result Adaptive-OPRO consistently outperforms fixed prompts in financial trading.
Study on limits of LLM-based multi-agent planning reliability.
problem Reliability limits of LLM-based multi-agent planning.
method Modeling LLM-based multi-agent architecture as a decision network, showing dominance by centralized Bayes decision maker.
result Optimizing multi-agent directed acyclic graphs under communication budget is equivalent to choosing a constrained experiment.
Bayesian principles improve agentic AI decision-making.
problem Decision-making under uncertainty in agentic AI systems.
method Bayesian decision theory applied to the orchestration layer of agentic AI.
result Bayesian principles enhance agentic AI's ability to make decisions under uncertainty.
Study uses LLMs for financial sentiment analysis without fine-tuning.
problem Lack of prescriptive knowledge to leverage generative models in FSA.
method Proposes a design framework with heterogeneous LLM agents based on Minsky's theory.
result Framework yields better accuracies, especially with substantial discussions.
New benchmarks focus on LLM risk in finance, not just accuracy.
problem Standard benchmarks ignore LLM financial risks, leading to unsafe deployment.
method Three-level agenda: model, workflow, and system stress-testing.
result Hidden weaknesses in LLMs are missed by traditional benchmarks.
LLMs in financial markets show diverse behaviors, from stable to speculative, challenging rational expectations.
problem Understanding the economic behaviors of LLMs in financial markets.
method Simulated financial market with 15 LLMs of varying sizes and capabilities.
result LLMs exhibit a spectrum of behaviors, including speculative bubbles, inconsistent with rational expectations.
Study examines if LLMs' trading styles match real market behavior.
problem Lack of behavioral consistency in LLMs' trading strategies.
method Year-long simulations with LLMs, operationalizing behavioral finance drivers, and comparing with financial theory.
result LLMs' strategy switching is only partially consistent with behavioral finance theories.
LLM agents discover cryptocurrency factors under reproducible constraints.
problem Flexibility of LLM agents in empirical discovery leads to uncontrolled search.
method Sequential hypothesis search with fixed data splits and portfolio tests.
result Ridge-combined portfolio achieves 44.55% annualized return in out-of-sample period.
The paper proposes a new method to evaluate LLM agent responses using ECDF clustering.
problem The standard evaluation of LLM agent responses via majority voting obscures response quality and distribution.
method The paper introduces a novel evaluation framework based on ECDF of cosine similarities and clustering of ECDFs using distances and k-medoids algorithm. result ECDF clustering reveals interpretable group structures in LLM responses, offering insights into agent settings.
Developing an AI economist agent using RAG, knowledge graphs, and LLMs for economic scenario analysis.
problem Economic scenario analysis using large language models and knowledge graphs.
method Proposing an RAG-based AI economist framework that utilizes knowledge graphs and LLMs.
result Improves economic coherence and traceability in generated reports.
LLMs add value in commodity portfolio construction when information set and implementation rules are held fixed.
problem Commodity portfolio construction
method Multi-Agent LLM Framework
result LLM strategies outperform Rule Agent in Sharpe terms
Unified LLM-agent with RL improves financial trading performance.
problem LLMs struggle with complex, multi-step financial tasks.
method Fusion of LLMs and gradient-based RL for policy optimization.
result Improves LLM performance in trading and other financial tasks.
LLMs can collude in market divisions, maximizing profits.
problem Strategic collusion of LLM agents in multi-commodity markets.
method Examined LLMs in Cournot competition frameworks, analyzing pricing and resource allocation strategies.
result LLMs can monopolize specific commodities without direct human input or explicit collusion commands.
Enhanced financial trading system using multi-agent LLMs with layered memory.
problem Inefficient prioritization of tasks in LLMs due to their memory processing.
method Introducing a multi-agent framework with layered memories and inter-agent debate.
result Superior automated trading accuracy and decision robustness.
The paper optimizes LLM inference systems through queueing theory.
problem Efficient LLM inference for AI agents under various routing topologies.
method Developed a fluid-limit framework for multi-class batched processing networks under K-FCFS scheduling.
result Proved that work-conserving scheduling algorithms maximize throughput for LLM inference.
Improved investment performance with fine-grained LLM tasks.
problem Abstract financial trading systems often overlook real-world workflow intricacies, leading to degraded performance.
method Proposes a multi-agent LLM trading framework that decomposes investment analysis into fine-grained tasks.
result Fine-grained task decomposition significantly improves risk-adjusted returns compared to coarse-grained designs.
PandaAI: A practical agent for neuro-symbolic data analysis and decision-making in finance
problem Sequential decision-making in finance
method Leveraging LLMs for market regime modeling and constrained alpha generation
result PandaAI achieves higher Rank IC and lower maximum drawdown
LLMs simulate financial markets, revealing consistent trading strategies and market dynamics.
problem Testing financial theories with AI trading agents.
method Simulated stock market with LLMs using a persistent order book and varied strategies.
result LLMs can simulate different trading strategies and market dynamics.
TRIBE model uses LLMs to simulate human trading behavior in bond markets.
problem Complexities in decentralized bond market transactions.
method Agent-based model augmented with LLMs to simulate human-like decision-making.
result Slight trade aversion in LLMs can lead to complete market collapse.
ADR helps LLMs find and use historical analogies for foresight analysis.
problem LLMs struggle to find relevant historical analogies due to surface-level matching.
method Proposes CANA framework with mechanism alignment and cross-analogy confirmation.
result CANA improves historical analogy generation by up to 10%.
HedgeAgents boosts financial trading with balanced strategies.
problem Inefficient trading strategies under rapid market changes.
method Integrates LLMs with multi-agent system for robust decision-making.
result 70% annualized return and 400% total return over 3 years.
ContestTrade uses competitive teams to improve LLM trading performance.
problem High sensitivity to market noise in LLM-based trading systems.
method Internal competitive mechanism, data and research teams, real-time evaluation.
result Significantly outperforms other systems across various metrics.
A blindfolded LLM trading framework validates market signals without ticker memorization.
problem Ensuring LLMs trade based on genuine market understanding, not memorized data.
method Anonymize tickers and company names, verify signals through reasoning embeddings, and use PPO-DSR policy.
result Achieved Sharpe ratio of 1.40 +/- 0.22 across 20 seeds, robust in volatile markets.
A trading system uses LLMs to adapt to volatile crypto markets.
problem Volatility and market sentiment in cryptocurrencies make traditional models ineffective.
method Specialized LLM agents for technical analysis, sentiment evaluation, and decision-making; verbal feedback for continuous improvement.
result Agents outperform buy-and-hold strategy with consistent gains across market phases.
TIM framework uses LLMs and domain experts to infer DeFi user transaction intents.
problem Challenges in understanding user intent in DeFi transactions due to complex interactions and opaque logs.
method TIM framework leverages a DeFi intent taxonomy, multi-agent LLM system, and a Meta-Level Planner.
result TIM significantly outperforms existing methods in inferring user transaction intents.
This paper combines LLMs with RL for better trading strategies.
problem Myopic behavior and opaque policies in RL for trading.
method LLMs generate strategic trading advice to guide RL agents.
result LLM-guided RL agents outperform unguided RL in return and risk metrics.
Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.
problem How do LLM-advisors perform in complex financial domains where domain expertise is crucial?
method Lab-based user study with 64 participants, focusing on three challenges: preference elicitation, personalized guidance, and relationship building.
result LLM-advisors can match human performance in preference elicitation but struggle with conflicting needs and trust issues.
LLM-based trading systems vary in execution realism and reproducibility.
problem LLM-based trading systems vary in execution realism and reproducibility.
method Reproducibility audit of 30 trade-relevant primary studies.
result LLM-based trading systems vary in execution realism and reproducibility.
Paper reviews the evolution of alpha from human insight to AI-powered systems.
problem Exceeding market benchmarks in finance.
method Five-stage taxonomy integrating representation learning, multimodal data fusion, and LLM agents.
result Unified framework for evaluating and developing next-gen alpha systems.
Transformer models can approximate smooth functions with prompts, enhancing LLMs' dynamic capabilities.
problem Lack of theoretical framework for prompt engineering in transformer models.
method Formal framework demonstrating transformer models can approximate β-times differentiable functions with prompts. result Transformer models can approximate β-times differentiable functions with arbitrary precision using appropriately structured prompts.