Study examines if LLMs' trading styles match real market behavior.
problem Lack of behavioral consistency in LLMs' trading strategies.
method Year-long simulations with LLMs, operationalizing behavioral finance drivers, and comparing with financial theory.
result LLMs' strategy switching is only partially consistent with behavioral finance theories.
Optimizes sparse mean-reverting portfolios for higher returns.
problem Finding optimal stock weights for mean-reverting portfolios.
method Transformed optimization problem into SDP, added constraints.
result Sparse mean-reverting portfolios provide higher returns with transaction costs.
This paper analyzes DRL strategies in finance, revealing unique trading patterns and performance differences.
problem Limited research on DRL behavior in finance applications.
method Analysis of trading behaviors and purchase diversity of DRL algorithms (A2C, PPO, SAC, DDPG, TD3).
result DRL algorithms exhibit distinct trading patterns and performance differences, with A2C outperforming others in terms of cumulative rewards.
Study long-term asset liquidation behavior with external flows.
problem Investigate optimal liquidation in presence of external flows.
method Convergence analysis of BSDEs for value function and strategy.
result Long-term liquidation may not occur due to external flows.
Strategy evaluation schemes are a crucial factor in any agent-based market model, as they determine the agents' strategy preferences and consequently their behavioral pattern. This study investigates how the strategy evaluation schemes adopted by agents affect their performance in conjunction with the market circumstan…
Study Figgie card game strategies using agent-based simulation.
problem Analyze strategies for Figgie card game and market behavior.
method Develop agent-based discrete-event market simulation to test strategies.
result Fundamentalist strategy is profit-maximizing in all tested combinations.
Investor selects portfolios based on news attention in a hidden Markov model.
problem Mean-variance portfolio selection in a dynamic attention context.
method Closed-loop equilibrium strategies via extended HJB equation and Markov chain approximation.
result Equilibrium strategies found through iterative algorithm and numerical examples.
Trading bubbles form when traders adapt to price mismatches.
problem Self-sustained price bubbles driven by adaptive trading behavior.
method Multi-agent model illustrating price bubble formation and statistical properties.
result Price bubbles can be driven by adaptive investment strategies.
Study finds users mostly use recent market and decision information to guess market direction.
problem Limited ability to model and predict human decision-making in stock markets.
method Used networks inference with stochastic block models (SBM) to find most predictive model of unobserved decisions.
result Users mostly use recent information to guess market direction, and their decision-making strategies are analogous to behaviors in other contexts.
ClusterLOB clusters market events to identify different trading behaviors.
problem Understanding market microstructure and participant behavior in financial markets.
method ClusterLOB uses K-means++ algorithm to cluster market events based on six time-dependent features.
result ClusterLOB identifies three distinct trading behaviors: directional, opportunistic, and market-making participants.
A simple trading model based on pair pattern strategy space with holding periods is proposed. Power-law behaviors are observed for the return variance σ2, the price impact H and the predictability K for both models with linear and square root impact functions. The sum of the traders' wealth displays a positive v…
A strategy to beat benchmarks by investing in heavily shorted but fundamentally sound securities.
problem Overcoming behavioral biases in investing, particularly the 'rebound effect'.
method Quantitative metrics, historical data, and securities lending modeling.
result The Bounce Basket strategy can outperform market returns during market downturns.
Proposes a robust Q-learning method to improve treatment strategy estimation.
problem Misspecification of working models in Q-learning leads to confounding and efficiency loss.
method Uses data-adaptive techniques to estimate nuisance parameters robustly.
result Asymptotic behavior of robust Q-learning estimators is studied and shown to be useful.
Behavior Transfer improves reinforcement learning by leveraging pre-trained policies.
problem Efficient transfer of knowledge in reinforcement learning.
method Behavior Transfer (BT) that uses pre-trained policies for exploration.
result BT combined with pre-training leads to better solutions than without pre-training.
We develop a behavioral asset pricing model in which agents trade in a market with information friction. Profit-maximizing agents switch between trading strategies in response to dynamic market conditions. Due to noisy private information about the fundamental value, the agents form different evaluations about heteroge…
In this paper, making use of recent statistical physics techniques and models, we address the specific role of randomness in financial markets, both at the micro and the macro level. In particular, we review some recent results obtained about the effectiveness of random strategies of investment, compared with some of t…
Improved genetic programming by optimizing mutation operators for continuous program search.
problem Small syntactic mutations in genetic programming can lead to unpredictable behavioral shifts.
method Learned a compact trading-strategy DSL, created a block-factorized embedding, and designed geometry-compiled mutation operators.
result Geometry-compiled mutation operators discover strong strategies using fewer evaluations and achieve higher Sharpe ratios.
Investigates optimal strategies for behavioral control problems with finite variation controls.
problem Behavioral singular stochastic control problems with finite variation controls.
method Abstract framework, applied to storage management and portfolio investment problems, using CPT preferences and Skorokhod representation theorem.
result Existence of optimal strategies for various goal functionals, including CPT preferences.
The study models mortgage prepayment risk, accounting for behavioral uncertainty, and provides replication strategies.
problem Modeling and replicating the prepayment option of mortgages with behavioral uncertainty.
method Modeling behavioral uncertainty as a non-hedgeable risk factor, proving its impact on exposure value, and using IRSs and swaptions for replication.
result Including behavioral uncertainty reduces the exposure's value, and swaptions are necessary for optimal replication.
We introduce a new approach for comparing reinforcement learning policies, using Wasserstein distances (WDs) in a newly defined latent behavioral space. We show that by utilizing the dual formulation of the WD, we can learn score functions over policy behaviors that can in turn be used to lead policy optimization towar…
Method learns behavioral states from wearable sensor data.
problem Understanding behavioral patterns from sensor data.
method Non-parametric Bayesian approach to model sensor data.
result Learned behavioral states cluster participants into meaningful groups and predict psychological states.
Local Bayesian optimization shows strong performance and converges well, contrary to folklore.
problem Understanding the behavior and convergence of local Bayesian optimization methods.
method Studied the behavior of local optimization strategies and rigorously analyzed a specific algorithm.
result Local Bayesian optimization algorithms converge well and perform strongly, contrary to the folklore.
Combining global and local explanations improves user understanding of RL agents.
problem Challenges in explaining agent behavior due to large state spaces and delayed rewards.
method Integrating strategy summaries with saliency maps to provide both global and local explanations.
result Summaries including important states significantly improve user understanding of RL agents.
We consider the problem of stopping a diffusion process with a payoff functional that renders the problem time-inconsistent. We study stopping decisions of naive agents who reoptimize continuously in time, as well as equilibrium strategies of sophisticated agents who anticipate but lack control over their future selves…
A competing market model with a polyvariant profit function that assumes "zeitnot" stock behavior of clients is formulated within the banking portfolio medium and then analyzed from the perspective of devising optimal strategies. An associated Markov process method for finding an optimal choice strategy for monovariant…
We consider the problem of portfolio optimization in the presence of market impact, and derive optimal liquidation strategies. We discuss in detail the problem of finding the optimal portfolio under Expected Shortfall (ES) in the case of linear market impact. We show that, once market impact is taken into account, a re…
Study examines how traders with asymmetric information and adaptive learning strategies affect market efficiency.
problem Effect of traders' strategic behavior on market efficiency and informational asymmetry.
method Examines a market with boundedly rational, asymmetrically informed traders using multiarmed bandit algorithms.
result Strategically acting traders can lead to more efficient markets than purely competitive ones under certain conditions.
Simulates DeLend Platform behavior to optimize operational parameters.
problem Optimizing the DeLend Platform's operational parameters.
method Agent-based simulations to connect and test different agent sets.
result Estimates how key variables respond to different policies.
The paper analyzes strategic irreversible investments with novel dynamic strategies.
problem Tradeoff between preemption incentives and option value of waiting in oligopolistic markets.
method Developed novel Markov perfect equilibrium to handle singular control of optimal investment.
result Simpler strategies lead to a 'preemption trap' with zero net present values.
We introduce a trade strategy representation theorem for performance measurement and portable alpha in high frequency trading, by embedding a robust trading algorithm that describe portfolio manager market timing behavior, in a canonical multifactor asset pricing model. First, we present a spectral test for market timi…
We present a simple one-parameter model for spatially localised evolving agents competing for spatially localised resources. The model considers selling agents able to evolve their pricing strategy in competition for a fixed market. Despite its simplicity, the model displays extraordinarily rich behavior. In addition t…
The paper analyzes reinsurance strategies in a competitive multi-agent system.
problem Strategic interactions and competitive behavior in multi-layer reinsurance chains.
method Stochastic differential games and non-zero-sum game models to characterize strategic interactions. Dynamic programming and game theory to derive equilibrium strategies.
result Intensified competition reduces safety loadings in reinsurance contracts.
We introduce a model of super-exponential financial bubbles with two assets (risky and risk-free), in which rational investors and noise traders co-exist. Rational investors form expectations on the return and risk of a risky asset and maximize their constant relative risk aversion expected utility with respect to thei…
Deep reinforcement learning (RL) methods generally engage in exploratory behavior through noise injection in the action space. An alternative is to add noise directly to the agent's parameters, which can lead to more consistent exploration and a richer set of behaviors. Methods such as evolutionary strategies use param…
BEACON optimizes discovery by efficiently finding novel behaviors.
problem Discovering diverse system behaviors without a scalar objective.
method Bayesian optimization inspired strategy using multi-output Gaussian processes.
result BEACON discovers broader sets of distinct behaviors than competing methods.
This paper develops a framework for efficient decision-making under time pressure.
problem Efficient decision-making under time pressure and subjective tradeoffs.
method Unified framework for evidence-based decision-making under time pressure.
result Ability to model and understand decision-making behavior under time constraints.
Triple-GAIL learns from multiple sources to improve imitation learning for complex behaviors.
problem Limited scalability of GAIL in real-world scenarios like autonomous vehicles.
method Integrates expert demonstrations and generated experiences with an auxiliary skill selector.
result Triple-GAIL outperforms state-of-the-art methods in learning complex behaviors.
We introduce a minimal Agent Based Model for financial markets to understand the nature and Self-Organization of the Stylized Facts. The model is minimal in the sense that we try to identify the essential ingredients to reproduce the main most important deviations of price time series from a Random Walk behavior. We fo…
Modeling decision-making dynamics in MDD patients using RL-HMM.
problem Characterize reward learning strategies in MDD patients.
method Proposed RL-HMM framework to analyze reward-based decision-making.
result MDD patients show reduced engagement in RL compared to healthy controls.
Reinforcement Learning AI commonly uses reward/penalty signals that are objective and explicit in an environment -- e.g. game score, completion time, etc. -- in order to learn the optimal strategy for task performance. However, Human-AI interaction for such AI agents should include additional reinforcement that is impl…
Empirical evidence supports new financial market definitions.
problem Investor risk attitudes in financial markets.
method Developed a new method to analyze risk attitudes.
result Risk-averse behavior in equity investors, risk-loving behavior in risk-free asset investors.
AI beats 95% of humans in Rock-Paper-Scissors.
problem Predicting and modeling human behavior in strategic games.
method Used Markov Models of varying memory lengths to compete against humans, introducing a 'focus length' parameter.
result Multi-AI strategy wins over 95% of human opponents in continuous 300-round games.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
Due to the popularity of the Internet and smart mobile devices, more and more financial transactions and activities have been digitalized. Compared to traditional financial fraud detection strategies using credit-related features, customers are generating a large amount of unstructured behavioral data every second. In …
Robinhood users react strongly to overnight price changes and big losers, trading quickly after extreme losses.
problem Understanding trading behavior of Robinhood users, especially in high-frequency trading scenarios.
method Analyzed intraday and overnight price changes, focusing on big losers and gainers.
result Robinhood users react more to overnight price changes and big losers, trading quickly after extreme losses.
Behavior modification improves prediction accuracy by nudging user behavior.
problem Improving prediction accuracy using behavior modification techniques.
method Combining prediction and behavior modification with reinforcement learning algorithms.
result Behavior modification can make predictions more certain but may not generalize.
We show that the statistics of spreads in real order books is characterized by an intrinsic asymmetry due to discreteness effects for even or odd values of the spread. An analysis of data from the NYSE order book points out that traders' strategies contribute to this asymmetry. We also investigate this phenomenon in th…
Study optimizes dividend payout strategies under fluctuating interest rates.
problem Maximizing dividends under stochastic interest rates with negative values.
method Analytical HJB approach and backward SDEs for analysis.
result Explicit optimal strategies found for both time-dependent and strategy-independent stopping times.