Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

74147221294 · May 202619922001200920172026
48 results for annualized percentage yield

PolyBench benchmarks LLMs on real market data, revealing significant performance gaps.

problem Benchmarking LLMs for real-world event prediction from live market signals.
method Multimodal benchmark derived from Polymarket, evaluating 7 LLMs under identical market states.
result Only two models achieve positive financial returns, highlighting the gap between fluency and probabilistic reasoning.

End-to-end framework optimizes financial metrics using neural networks.

problem Difficult portfolio optimization in financial markets due to non-stationarity and high costs.
method Directly optimizes differentiable financial metrics via neural networks, incorporating realistic costs and rebalancing.
result Best model achieves +7.86% total return, outperforming S&P 500 by 12.38 percentage points.

The p-index improves investment performance for NYSE stocks but not for SSE stocks.

problem Improving investment performance for stocks using the p-index.
method Comparing different p-ratio strategies and empirical efficient frontiers for SSE and NYSE stocks.
result The p-index enhances investment performance for NYSE stocks but not for SSE stocks.

This paper derives a robust on-line equity trading algorithm that achieves the greatest possible percentage of the final wealth of the best pairs rebalancing rule in hindsight. A pairs rebalancing rule chooses some pair of stocks in the market and then perpetually executes rebalancing trades so as to maintain a target …

2018-10-04abs ↗pdf ↗

Study shows how business cycle affects dividend payout based on managerial stock incentives.

problem Impact of managerial stock incentives on dividend payout policy during business cycles.
method Using S&P 1500 companies data from 2000-2018, analyzing full sample and recession periods.
result Negative relationship between managerial stock options and dividend payouts, significant for medium-sized companies.

Bayesian consensus improves accuracy of forecasts from miscalibrated sources.

problem Aggregating predictions from miscalibrated and noisy sources.
method Bayesian approach to adjust for bias and noise, using hierarchical models.
result Bayesian consensus estimator is unbiased and more efficient than alternatives.

SAGA predicts multi-year earnings with adaptive intervals, improving forecast accuracy.

problem Forecasting long-range nonlinear structure in lifetime earnings.
method Decoder-only transformer for irregular tabular sequences, split conformal calibration.
result Significant improvement in forecast accuracy compared to existing methods.

A3T-GCN model forecasts FTSE100 stock prices using technical indicators and financial ratios.

problem Forecasting closing stock prices of FTSE100 constituents.
method Hybrid A3T-GCN architecture using technical indicators, financial ratios, and sector correlations.
result A3T-GCN model improves prediction accuracy with annualized log-returns and shorter sequence lengths.

Paper compares ETF and futures carry rates in segmented Bitcoin markets.

problem Limitations in cross-margining between spot Bitcoin and CME futures.
method Estimates carry rates from IBIT options and CME futures, uses put-call parity and daily ETF holdings.
result Mean and median wedge in carry rates is 2.58 and 2.52 percent, respectively.

Two machine learning models detect anomalies in ER claims, saving up to 40% in improper payments.

problem Improper health insurance payments from fraud and upcoding.
method Two machine learning models: an upcoding model based on severity code distributions and a random forest model for claim sorting.
result Random forest model saved 12% to 40% in improper payments compared to a baseline approach.

The paper models US inflation and hyperinflation using monetary and GDP data.

problem Understanding and predicting inflation and hyperinflation.
method Developed economic models to predict US CPI growth based on BMS, GDP, and savings.
result An exact relationship between CPI growth and BMS growth minus GDP and savings growth was found, with a residual term.

Study shows houses appreciated more during pandemic due to speculation, not just price uncertainty.

problem Impact of COVID-19 on house prices and speculation.
method Quasi-experimental design, unit-level matching, multivariate difference-in-difference regression.
result Properties listed for sale appreciated an additional 1% per month after pandemic onset, with an excess annual growth of 12.7 percentage points.

Framework integrates financial and annual report data for better corporate credit ratings.

problem Lack of insights from non-financial data in credit rating models.
method Uses FinBERT to extract features from annual reports and combines them with financial data.
result Improves credit rating accuracy by 8-12%.

Study shows survivorship bias inflates returns in India's small-cap index.

problem Survivorship bias in emerging market small-cap indices.
method Reconstructing historical index composition through market capitalization ranking and comparing equal-weight portfolios of current constituents versus all historical members.
result Survivor-only backtesting overstates returns by 4.94 percentage points and Sharpe ratios by 0.097.

The study extends SPT to account for real-world transaction costs, improving portfolio performance.

problem Real-world transaction costs affect portfolio performance, especially during market stress.
method Developed a continuous-time model with stochastic transaction costs and derived lower bounds for cost-adjusted wealth.
result Functionally generated portfolios can still achieve relative arbitrage after accounting for transaction costs.

Maximizes probability of completing investment schedules with optimal portfolio weights.

problem Optimizing probability of completing investment schedules with optimal portfolio weights.
method Computing maximum probability and optimal portfolio weight functions for various rebalancing schedules.
result Noticeable improvements in probability to complete schedules with optimal portfolio weights.

Classification outperforms regression in portfolio construction, yielding higher Sharpe ratios.

problem Determining which machine learning approach (classification vs. regression) is more effective for portfolio construction.
method Used stacking ensemble of gradient boosted tree, random forest, and neural network models.
result Classification yields higher Sharpe ratios and economically significant alphas compared to regression.

Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…

2011-02-17abs ↗pdf ↗

Paper uses LLMs to analyze annual reports for stock investment, improving efficiency.

problem Manual analysis of annual reports is time-consuming and requires expertise.
method Leverages Large Language Models to extract and analyze annual reports.
result Machine Learning model trained on LLM outputs outperforms S&P500 returns.

Weak predictability of stock price movement 2 days after annual report disclosure.

problem Predicting stock price movement after annual report disclosure.
method Used various models including decision tree, logistic regression, random forest, neural network, prototypical networks; used financial indicators from EastMoney.
result Maximum accuracy and precision of stock price movement prediction is around 59.6% and 0.56 respectively, with random forest performing best.

Data analytics and machine learning techniques are being rapidly adopted into the power system, including power system control as well as electricity market design. In this paper, from an adversarial machine learning point of view, we examine the vulnerability of data-driven electricity market design. More precisely, w…

2019-11-18abs ↗pdf ↗

We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression. We show that universal consistency of Empirical Risk Minimiza…

2015-06-12abs ↗pdf ↗

The study forecasts portfolio volatility using cointegrated asset dynamics.

problem Forecasting volatility in portfolios with high accuracy.
method Developed HVR/DVR ratios and used Vector Error Correction Model (VECM) to forecast volatility.
result VECM forecasts of portfolio volatility have lower MAPE than covariance-based forecasts.

New simulations advise caution in choosing principal components for multivariate functional data.

problem Inaccurate selection of principal components in multivariate functional data.
method Extensive simulations investigating the reliability of percentage of variance explained thresholds.
result Conventional threshold methods may fail to accurately explain overall variance in multivariate functional data.

Study estimates Medallion's compounded return before fees at 31.8%.

problem Incorrectly using yearly returns for compounding leads to overestimation of fund performance.
method Used fund sizes and trading profits to estimate compounded return; used manager's wealth as proxy for Simons.
result Annualized compounded return of Medallion before fees is likely under 35%

This paper explores leverage staking with stETH, revealing high returns but also significant risks.

problem Leverage staking introduces risks through intensified selling pressure and cascading liquidations.
method Formal framework for leverage staking, stress tests under extreme conditions of stETH devaluation.
result Leverage staking amplifies risks, leading to intensified selling pressure and price declines.

We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We prove the existence of an optimal MAPE model and we show the universal consistency of Empirical Risk Minimization based on the MAPE. We also show that finding the best model under…

2016-05-09abs ↗pdf ↗

Duel-Evolve uses LLM self-preferences for test-time optimization of discrete outputs.

problem Optimizing LLM outputs at test time with limited or unreliable scalar rewards.
method Duel-Evolve uses pairwise comparisons from the LLM to guide optimization, aggregating them via a Bayesian Bradley-Terry model.
result Achieves significant improvement over existing methods in accuracy.

We study T. Cover's rebalancing option (Ordentlich and Cover 1998) under discrete hindsight optimization in continuous time. The payoff in question is equal to the final wealth that would have accrued to a $\$1$ deposit into the best of some finite set of (perhaps levered) rebalancing rules determined in hindsight. A r…

2019-03-03abs ↗pdf ↗

We uncover a large and significant low-minus-high rank effect for commodities across two centuries. There is nothing anomalous about this anomaly, nor is it clear how it can be arbitraged away. Using nonparametric econometric methods, we demonstrate that such a rank effect is a necessary consequence of a stationary rel…

2016-07-26abs ↗pdf ↗

Study predicts doubling of U.S. maize insurance claims due to climate change.

problem Climate change increases U.S. maize loss probability, impacting insurance claims.
method Neural Network Monte Carlo simulations to predict crop loss metrics.
result Doubling of annual probability of maize Yield Protection insurance claims by mid-century.

New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.

problem Crowding in index portfolios during reconstitution events.
method Developed a Python package for accurate index reconstruction using CRSP US Stock data.
result Annual Russell 3000 portfolios are more crowded than quarterly ones, suggesting lower transaction costs.