PolyBench benchmarks LLMs on real market data, revealing significant performance gaps.
problem Benchmarking LLMs for real-world event prediction from live market signals.
method Multimodal benchmark derived from Polymarket, evaluating 7 LLMs under identical market states.
result Only two models achieve positive financial returns, highlighting the gap between fluency and probabilistic reasoning.
Model predicts S&P 500 IT sector index prices with high accuracy.
problem Predicting S&P 500 IT sector index prices accurately.
method Non-linear model using financial and economic indicators.
result Predictive accuracy of 99.4% for S&P 500 IT sector index.
End-to-end framework optimizes financial metrics using neural networks.
problem Difficult portfolio optimization in financial markets due to non-stationarity and high costs.
method Directly optimizes differentiable financial metrics via neural networks, incorporating realistic costs and rebalancing.
result Best model achieves +7.86% total return, outperforming S&P 500 by 12.38 percentage points.
Repo dealers' market power affects bond prices by up to 2 percentage points.
problem Market power of repo dealers impacts bond prices and liquidity.
method Proprietary data on repo and reverse-repo trades analyzed.
result Market power of repo dealers accounts for 0.5-1.3 percentage points of bond yield deviation.
The p-index improves investment performance for NYSE stocks but not for SSE stocks.
problem Improving investment performance for stocks using the p-index.
method Comparing different p-ratio strategies and empirical efficient frontiers for SSE and NYSE stocks.
result The p-index enhances investment performance for NYSE stocks but not for SSE stocks.
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach allows a provision for reduction of capital as a result of insurance mitigation of up to 20%. This paper studies the behaviour of different insurance policies in the context of capital reduction for a range of possible extreme los…
This paper derives a robust on-line equity trading algorithm that achieves the greatest possible percentage of the final wealth of the best pairs rebalancing rule in hindsight. A pairs rebalancing rule chooses some pair of stocks in the market and then perpetually executes rebalancing trades so as to maintain a target …
Study shows how business cycle affects dividend payout based on managerial stock incentives.
problem Impact of managerial stock incentives on dividend payout policy during business cycles.
method Using S&P 1500 companies data from 2000-2018, analyzing full sample and recession periods.
result Negative relationship between managerial stock options and dividend payouts, significant for medium-sized companies.
A new model detects financial bubbles with high accuracy.
problem Quantifying and detecting financial bubbles.
method Hyped Log-Periodic Power Law Model (HLPPL) with sentiment scores and hype index.
result Achieved an average annualized return of 34.13% during backtesting.
Bayesian consensus improves accuracy of forecasts from miscalibrated sources.
problem Aggregating predictions from miscalibrated and noisy sources.
method Bayesian approach to adjust for bias and noise, using hierarchical models.
result Bayesian consensus estimator is unbiased and more efficient than alternatives.
SAGA predicts multi-year earnings with adaptive intervals, improving forecast accuracy.
problem Forecasting long-range nonlinear structure in lifetime earnings.
method Decoder-only transformer for irregular tabular sequences, split conformal calibration.
result Significant improvement in forecast accuracy compared to existing methods.
A3T-GCN model forecasts FTSE100 stock prices using technical indicators and financial ratios.
problem Forecasting closing stock prices of FTSE100 constituents.
method Hybrid A3T-GCN architecture using technical indicators, financial ratios, and sector correlations.
result A3T-GCN model improves prediction accuracy with annualized log-returns and shorter sequence lengths.
Paper compares ETF and futures carry rates in segmented Bitcoin markets.
problem Limitations in cross-margining between spot Bitcoin and CME futures.
method Estimates carry rates from IBIT options and CME futures, uses put-call parity and daily ETF holdings.
result Mean and median wedge in carry rates is 2.58 and 2.52 percent, respectively.
Two machine learning models detect anomalies in ER claims, saving up to 40% in improper payments.
problem Improper health insurance payments from fraud and upcoding.
method Two machine learning models: an upcoding model based on severity code distributions and a random forest model for claim sorting.
result Random forest model saved 12% to 40% in improper payments compared to a baseline approach.
Prediction markets and crypto options show persistent pricing gaps.
problem Comparing prediction markets and crypto options for identical payoffs.
method Comparing Polymarket Yes prices with Binance call option prices.
result Mean pricing gap of 5.6 percentage points across 214 hourly observations.
The paper models US inflation and hyperinflation using monetary and GDP data.
problem Understanding and predicting inflation and hyperinflation.
method Developed economic models to predict US CPI growth based on BMS, GDP, and savings.
result An exact relationship between CPI growth and BMS growth minus GDP and savings growth was found, with a residual term.
Study shows houses appreciated more during pandemic due to speculation, not just price uncertainty.
problem Impact of COVID-19 on house prices and speculation.
method Quasi-experimental design, unit-level matching, multivariate difference-in-difference regression.
result Properties listed for sale appreciated an additional 1% per month after pandemic onset, with an excess annual growth of 12.7 percentage points.
Framework integrates financial and annual report data for better corporate credit ratings.
problem Lack of insights from non-financial data in credit rating models.
method Uses FinBERT to extract features from annual reports and combines them with financial data.
result Improves credit rating accuracy by 8-12%.
Study shows survivorship bias inflates returns in India's small-cap index.
problem Survivorship bias in emerging market small-cap indices.
method Reconstructing historical index composition through market capitalization ranking and comparing equal-weight portfolios of current constituents versus all historical members.
result Survivor-only backtesting overstates returns by 4.94 percentage points and Sharpe ratios by 0.097.
The study extends SPT to account for real-world transaction costs, improving portfolio performance.
problem Real-world transaction costs affect portfolio performance, especially during market stress.
method Developed a continuous-time model with stochastic transaction costs and derived lower bounds for cost-adjusted wealth.
result Functionally generated portfolios can still achieve relative arbitrage after accounting for transaction costs.
Maximizes probability of completing investment schedules with optimal portfolio weights.
problem Optimizing probability of completing investment schedules with optimal portfolio weights.
method Computing maximum probability and optimal portfolio weight functions for various rebalancing schedules.
result Noticeable improvements in probability to complete schedules with optimal portfolio weights.
Classification outperforms regression in portfolio construction, yielding higher Sharpe ratios.
problem Determining which machine learning approach (classification vs. regression) is more effective for portfolio construction.
method Used stacking ensemble of gradient boosted tree, random forest, and neural network models.
result Classification yields higher Sharpe ratios and economically significant alphas compared to regression.
In this paper we implement a Local Linear Regression Ensemble Committee (LOLREC) to predict 1-day-ahead returns of 453 assets form the S&P500. The estimates and the historical returns of the committees are used to compute the weights of the portfolio from the 453 stock. The proposed method outperforms benchmark portfol…
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…
We analyze annual revenues and earnings data for the 500 largest-revenue U.S. companies during the period 1954-2007. We find that mean year profits are proportional to mean year revenues, exception made for few anomalous years, from which we postulate a linear relation between company expected mean profit and revenue. …
Paper uses LLMs to analyze annual reports for stock investment, improving efficiency.
problem Manual analysis of annual reports is time-consuming and requires expertise.
method Leverages Large Language Models to extract and analyze annual reports.
result Machine Learning model trained on LLM outputs outperforms S&P500 returns.
In this paper we study a class of insurance products where the policy holder has the option to insure k of its annual Operational Risk losses in a horizon of T years. This involves a choice of k out of T years in which to apply the insurance policy coverage by making claims against losses in the given year. The…
A simple formula approximates AUM fees' cumulative costs.
problem Estimating the total cost of AUM fees over time.
method Intuitive explanation and analytical derivation of a formula.
result Investments lose almost Nε% of their value over N years with an annual fee of ε%.
A reliable and accurate forecasting model for crop yields is of crucial importance for efficient decision-making process in the agricultural sector. However, due to weather extremes and uncertainties, most forecasting models for crop yield are not reliable and accurate. For measuring the uncertainty and obtaining furth…
Tether's dominance in U.S. Treasury bills lowers bond yields by 24 basis points.
problem Impact of Tether's market share on U.S. Treasury bill yields.
method Baseline semi-log time trend model and threshold regression analysis.
result Tether's market share reduces 1-month yields by 24 basis points.
Weak predictability of stock price movement 2 days after annual report disclosure.
problem Predicting stock price movement after annual report disclosure.
method Used various models including decision tree, logistic regression, random forest, neural network, prototypical networks; used financial indicators from EastMoney.
result Maximum accuracy and precision of stock price movement prediction is around 59.6% and 0.56 respectively, with random forest performing best.
Data analytics and machine learning techniques are being rapidly adopted into the power system, including power system control as well as electricity market design. In this paper, from an adversarial machine learning point of view, we examine the vulnerability of data-driven electricity market design. More precisely, w…
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We show that universal consistency of Empirical Risk Minimiza…
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We also show that, under some asumptions, universal consisten…
The study forecasts portfolio volatility using cointegrated asset dynamics.
problem Forecasting volatility in portfolios with high accuracy.
method Developed HVR/DVR ratios and used Vector Error Correction Model (VECM) to forecast volatility.
result VECM forecasts of portfolio volatility have lower MAPE than covariance-based forecasts.
New simulations advise caution in choosing principal components for multivariate functional data.
problem Inaccurate selection of principal components in multivariate functional data.
method Extensive simulations investigating the reliability of percentage of variance explained thresholds.
result Conventional threshold methods may fail to accurately explain overall variance in multivariate functional data.
Study estimates Medallion's compounded return before fees at 31.8%.
problem Incorrectly using yearly returns for compounding leads to overestimation of fund performance.
method Used fund sizes and trading profits to estimate compounded return; used manager's wealth as proxy for Simons.
result Annualized compounded return of Medallion before fees is likely under 35%
This paper explores leverage staking with stETH, revealing high returns but also significant risks.
problem Leverage staking introduces risks through intensified selling pressure and cascading liquidations.
method Formal framework for leverage staking, stress tests under extreme conditions of stETH devaluation.
result Leverage staking amplifies risks, leading to intensified selling pressure and price declines.
Study integrates ESG factors into home price predictions for U.S. cities.
problem Predicting average annual home prices using ESG factors.
method Used P-spline GAM and GLM models, transformed time series data.
result ESG factors influence home prices differently by city.
This paper proposes a paradigm shift in the valuation of long term annuities, away from classical no-arbitrage valuation towards valuation under the real world probability measure. Furthermore, we apply this valuation method to two examples of annuity products, one having annual payments linked to a mortality index and…
We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We prove the existence of an optimal MAPE model and we show the universal consistency of Empirical Risk Minimization based on the MAPE. We also show that finding the best model under…
Duel-Evolve uses LLM self-preferences for test-time optimization of discrete outputs.
problem Optimizing LLM outputs at test time with limited or unreliable scalar rewards.
method Duel-Evolve uses pairwise comparisons from the LLM to guide optimization, aggregating them via a Bayesian Bradley-Terry model.
result Achieves significant improvement over existing methods in accuracy.
This paper uses alternative data to forecast Japanese real estate performance.
problem Accurate rent and price forecasting in Japanese real estate markets.
method Created a comprehensive house price index using over 5 million transactions and economic factors.
result Alternative data variables can forecast real estate performance effectively.
This dataset contains the annual aggregated income taxes of all the Italian municipalities over the years 2007-2011. Data are clustered over the Italian regions and provinces. The source of the data is the Italian Ministry of Economics and Finance. The administrative variations in Italy over the quinquennium have been …
We study T. Cover's rebalancing option (Ordentlich and Cover 1998) under discrete hindsight optimization in continuous time. The payoff in question is equal to the final wealth that would have accrued to a $\$1$ deposit into the best of some finite set of (perhaps levered) rebalancing rules determined in hindsight. A r…
We uncover a large and significant low-minus-high rank effect for commodities across two centuries. There is nothing anomalous about this anomaly, nor is it clear how it can be arbitraged away. Using nonparametric econometric methods, we demonstrate that such a rank effect is a necessary consequence of a stationary rel…
Study predicts doubling of U.S. maize insurance claims due to climate change.
problem Climate change increases U.S. maize loss probability, impacting insurance claims.
method Neural Network Monte Carlo simulations to predict crop loss metrics.
result Doubling of annual probability of maize Yield Protection insurance claims by mid-century.
New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.
problem Crowding in index portfolios during reconstitution events.
method Developed a Python package for accurate index reconstruction using CRSP US Stock data.
result Annual Russell 3000 portfolios are more crowded than quarterly ones, suggesting lower transaction costs.