Calibrating a trading rule using a historical simulation (also called backtest) contributes to backtest overfitting, which in turn leads to underperformance. In this paper we propose a procedure for determining the optimal trading rule (OTR) without running alternative model configurations through a backtest engine. We…
This work presents a theoretical and empirical evaluation of Anderson-Darling test when the sample size is limited. The test can be applied in order to backtest the risk factors dynamics in the context of Counterparty Credit Risk modelling. We show the limits of such test when backtesting the distributions of an intere…
We study a class of backtests for forecast distributions in which the test statistic depends on a spectral transformation that weights exceedance events by a function of the modeled probability level. The weighting scheme is specified by a kernel measure which makes explicit the user's priorities for model performance.…
This paper introduces novel backtests for the risk measure Expected Shortfall (ES) following the testing idea of Mincer and Zarnowitz (1969). Estimating a regression framework for the ES stand-alone is infeasible, and thus, our tests are based on a joint regression for the Value at Risk and the ES, which allows for dif…
Machine learning creates non-factor covariance matrices for risk models.
problem Creating robust risk models for financial portfolios.
method Developed an explicit algorithm and source code for machine learning risk models.
result Machine learning models outperform traditional risk models in empirical backtests.
This paper proposes a new method to prevent backtesting overfitting in trading strategies.
problem Preventing misleading results in backtesting of trading strategies.
method Covariance-Penalty Correction approach to reduce risk metrics based on the number of parameters and data used.
result Covariance-Penalties are effective in avoiding backtesting overfitting, with Total Least Squares outperforming Ordinary Least Squares.
Expected Shortfall (ES) has been widely accepted as a risk measure that is conceptually superior to Value-at-Risk (VaR). At the same time, however, it has been criticised for issues relating to backtesting. In particular, ES has been found not to be elicitable which means that backtesting for ES is less straightforward…
A new stock selection strategy uses combined machine learning with dynamic weighting methods.
problem Improving stock selection accuracy and performance.
method Combined machine learning algorithms with static and dynamic weighting methods.
result IC-based dynamic weighting outperforms static evaluation metrics in backtested returns and predictive performance.
This paper investigates bias in resampled backtests for financial portfolios, finding it often negligible.
problem Bias in resampled backtests for financial portfolio evaluation.
method Investigation of bias in rolling-window mean-variance portfolios using resampling techniques.
result The bias in Sharpe Ratio estimates from IID resampling is often a fraction of estimation noise, making it tolerable.
A new method for backtesting ES forecasts in banking.
problem Designing a model-free backtesting procedure for Expected Shortfall forecasts.
method Use e-values and e-processes to introduce backtest e-statistics for VaR and ES.
result The proposed method can be applied to various risk measures and statistical quantities.
This research improves forecasting and testing of risk contributions using Expected Shortfall.
problem Improving risk allocation and testing methods for regulatory standards.
method Developed a comprehensive framework for backtesting and forecasting Expected Shortfall contributions.
result Proposed a novel semiparametric model for forecasting dynamic Expected Shortfall contributions.
New method tests risk measures for various distortions.
problem Testing risk measures for different distortions.
method Stratification and randomization of risk levels.
result Method performs well in numerical case studies.
Study compares quantum and classical ML in crypto trading, finding hybrid models outperform.
problem Comparing quantum and classical machine learning in crypto trading strategies.
method Backtesting 10 models across multiple crypto assets using classical ML, quantum ML, hybrid models, and transformer models.
result Hybrid quantum models achieve superior performance with 13.99% return and 1.76 Sharpe ratio.
This paper defines systematic value investing as an empirical optimization problem. Predictive modeling is introduced as a systematic value investing methodology with dynamic and optimization features. A predictive modeling process is demonstrated using financial metrics from Gray & Carlisle and Buffett & Clark. A 31-y…
A new property fixes look-ahead bias in backtesting and trading pipelines.
problem Fixing look-ahead bias in backtesting and trading pipelines.
method Developed a pipeline calculus separating availability from reference time, and a type-and-effect system for the value-independent fragment.
result The check scales linearly and catches all leaks, including those missed by differential and tiling detectors.
The paper proposes a method to discount backtest PnLs due to in-sample overfitting.
problem In-sample overfitting in backtest-based investment strategies.
method A simple framework to model and quantify in-sample PnL overfitting.
result Computes the appropriate discount factor for PnLs of in-sample investment strategies.
Conditional forecasts of risk measures play an important role in internal risk management of financial institutions as well as in regulatory capital calculations. In order to assess forecasting performance of a risk measurement procedure, risk measure forecasts are compared to the realized financial losses over a perio…
Backtesting framework for CLMMs on Uniswap V3 reduces reward estimation error.
problem Estimating rewards for CLMMs in Uniswap V3 liquidity pools.
method Parametric model for liquidity distribution, historical data analysis.
result Error in reward estimation less than 1% for each pool.
In recent years several trading platforms appeared which provide a backtest engine to calculate historic performance of self designed trading strategies on underlying candle data. The construction of a correct working backtest engine is, however, a subtle task as shown by Maier-Paape and Platen (cf. arXiv:1412.5558 [q-…
Backtests of structured strategies lose much of their predictive power in live trading.
problem Uncertainty in how marketed backtests predict live performance of structured strategies.
method Analysis of 1,726 structured strategies from ten global institutions.
result Raw backtests have limited portability into live trading and deteriorate sharply.
The paper tackles backtest overfitting in cryptocurrency trading using deep reinforcement learning.
problem Backtest overfitting in deep reinforcement learning for cryptocurrency trading.
method Formulated hypothesis test for overfitting detection, trained agents, estimated overfitting probability, and rejected overfitted agents.
result Less overfitted deep reinforcement learning agents outperformed more overfitted agents and market benchmarks.
New method allows backtesting of systemic risk forecasts.
problem Systemic risk measures are not elitable and identifiable, making backtesting impossible.
method Introduces multi-objective elicitability and Diebold--Mariano type tests.
result Proposes a traffic-light approach for backtesting.
A new method tests Expected Shortfall by analyzing both duration and severity of VaR violations.
problem Lack of separate testing for frequency and severity in ES backtesting.
method Uses bivariate orthogonal polynomials to derive moment conditions for durations and severities.
result Proposes a Wald test for identifying mis-specified components in ES models.
The paper examines sizing strategies for algorithmic trading in volatile markets.
problem High volatility creates challenges for algorithmic traders.
method Investigates different sizing models and backtesting techniques for financial trading.
result Sizing models can lower Value at Risk (VaR) during crisis events.
GAS models have been recently proposed in time-series econometrics as valuable tools for signal extraction and prediction. This paper details how financial risk managers can use GAS models for Value-at-Risk (VaR) prediction using the novel GAS package for R. Details and code snippets for prediction, comparison and back…
The paper uses machine learning to simulate financial markets and improve trading strategy backtesting.
problem Improving risk management of quantitative investment strategies.
method Simulates financial markets using Boltzmann Machines and Generative Adversarial Networks to preserve asset return distributions and dependencies.
result Developed a framework to estimate backtest statistics more accurately.
In this paper we try to design the necessary calculation needed for backtesting trading systems when only candle chart data are available. We lay particular emphasis on situations which are not or not uniquely decidable and give possible strategies to handle such situations.
We propose a new backtesting framework for Expected Shortfall that could be used by the regulator. Instead of looking at the estimated capital reserve and the realised cash-flow separately, one could bind them into the secured position, for which risk measurement is much easier. Using this simple concept combined with …
Using non-linear machine learning methods and a proper backtest procedure, we critically examine the claim that Google Trends can predict future price returns. We first review the many potential biases that may influence backtests with this kind of data positively, the choice of keywords being by far the greatest culpr…
Improved stock return prediction model handles noise and non-stationarity.
problem Predicting stock returns with robustness to noise and non-stationarity.
method Extended AROW algorithm to handle synchronous mini-batch updates and applied it to stock return prediction.
result The new model outperforms classical approaches in backtesting on S\&P500 stocks.
We propose factor models for the cross-section of daily cryptoasset returns and provide source code for data downloads, computing risk factors and backtesting them out-of-sample. In "cryptoassets" we include all cryptocurrencies and a host of various other digital assets (coins and tokens) for which exchange market dat…
The paper estimates CoVaR with various models for financial risk analysis.
problem Estimating conditional value-at-risk with financial time series data.
method Fitting multivariate parametric models and copula functions to capture stylized facts of equity returns.
result Backtesting shows that certain models provide better risk estimates than others.
In this note, we comment on the relevance of elicitability for backtesting risk measure estimates. In particular, we propose the use of Diebold-Mariano tests, and show how they can be implemented for Expected Shortfall (ES), based on the recent result of Fissler and Ziegel (2015) that ES is jointly elicitable with Valu…
Benchmark detects decision-time leakage in financial backtests.
problem Detecting decision-time leakage in financial machine-learning backtests.
method Toggles one evaluation convention at a time around a clean t+1-open reference, holding other factors fixed. result Inflation is highly selective, affecting specific features and execution methods.
New method corrects risk estimation bias, improving backtesting results.
problem Underestimation of risk by existing methods, especially in small samples.
method Proposes a new algorithm for bias correction using generalized Pareto distributions.
result The new algorithm leads to improved efficiency in estimating risk with heavy tails or heteroscedasticity.
New metrics quantify implementation risk in portfolio backtesting, revealing systematic differences in engine implementations.
problem Systematic divergence in backtested portfolio metrics due to differences in engine implementations.
method Formalized implementation risk, proposed four metrics, executed 15 strategies through five engines, analyzed source-code defects.
result Implementation risk introduces measurable ambiguity in performance attribution, but does not alter investment decisions.
Credibility theory provides tools to obtain better estimates by combining individual data with sample information. We apply the Credibility theory to a Uniform distribution that is used in testing the reliability of forecasting an interest rate for long term horizons. Such empirical exercise is asked by Regulators (CRR…
The paper optimizes portfolios using clustering and Sharpe ratio-based optimization.
problem Optimizing portfolio performance in financial modeling.
method Combines K-Means clustering for asset segmentation and Sharpe ratio-based optimization.
result Optimized portfolios outperform traditional equal-weighted benchmarks.
A new risk measure, the lambda value at risk (Lambda VaR), has been recently proposed from a theoretical point of view as a generalization of the value at risk (VaR). The Lambda VaR appears attractive for its potential ability to solve several problems of the VaR. In this paper we propose three nonparametric backtestin…
AutoQuant addresses cryptocurrency backtesting fragility by modeling execution costs and improving strategy selection.
problem Fragile backtests of cryptocurrency perpetual futures ignoring microstructure frictions and execution costs.
method Execution-centric framework with Bayesian optimization, double screening, and strict T+1 semantics.
result Fee-only and zero-cost backtests overestimate returns, highlighting the importance of modeling execution costs.
Robust forecast framework reduces distribution error by 63%.
problem Accurate distribution forecast for planning decisions.
method Backtest-based bootstrap and adaptive residual selection.
result Reduces Absolute Coverage Error by more than 63%.
AlphaEval evaluates alpha mining models efficiently and comprehensively.
problem Lack of systematic evaluation for alpha mining models.
method Unified, parallelizable evaluation framework assessing predictive power, stability, robustness, financial logic, and diversity.
result AlphaEval achieves evaluation consistency comparable to comprehensive backtesting, providing more comprehensive insights and higher efficiency.
Cryptocurrency patterns stable across market caps, validated by microstructure theory.
problem Stable patterns in cryptocurrency microstructure across different market caps.
method Unified CatBoost modeling pipeline with time-series cross validation, validated by backtests.
result Feature rankings and partial effects are stable across assets despite heterogeneous liquidity and volatility.
AlphaLogics mines market logic to generate interpretable alpha factors.
problem Complex, opaque alpha factors from factor mining overlook market logic.
method Market Logic Mining, Factor Generation and Optimization, Market Logic Generation and Optimization.
result AlphaLogics improves predictive metrics and risk-adjusted returns over baselines.
Paper proposes SPO paradigm for better portfolio optimization in real markets.
problem Real-world trading frictions and constraints affect portfolio optimization quality.
method SPO paradigm with decision-focused training using surrogate loss and linear predictors.
result Decision-focused training improves risk-adjusted performance and robustness.
This paper evaluates LLMs for technical market analysis, finding GPT-4 Turbo and FinGPT outperform passive benchmarks.
problem Evaluating LLMs for technical market analysis in financial markets.
method Structured evaluation of five LLMs (GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, FinGPT) on four tasks: candlestick pattern recognition, directional signal generation, backtesting, and financial report comprehension.
result GPT-4 Turbo and FinGPT outperform passive benchmarks in simulated backtesting, with GPT-4 Turbo achieving the highest annualized return and Sharpe ratio.
Anonymizing company names in financial news improves trading performance, contrary to initial expectations.
problem Look-ahead and distraction biases in sentiment analysis of financial news.
method Investigated trading strategies based on original and anonymized headlines, comparing performance.
result Anonymized headlines outperform original in-sample, suggesting distraction effect is stronger.
New algorithm corrects risk estimation bias for heavy-tailed data.
problem Underestimation of risk in banking and insurance due to bias in estimation procedures.
method Proposes a new algorithm for bias correction and applies it to generalized Pareto distributions.
result The algorithm leads to more accurate risk estimation, especially in heavy-tailed data.