Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

73146218291 · Jun 202019922001200920182026
48 results for Historical Back Testing

Paper introduces a new risk measure for multivariate residual estimation.

problem Quantifying residual estimation risk in complex financial models.
method Developed a multivariate framework for residual estimation risk, defined using various risk measures, and proposed a back-testing criterion.
result Demonstrated the effectiveness of the new measure through back-testing on retail credit portfolios.

Proposes dynamic borrowing method for historical data in clinical trials.

problem Insufficient statistical power in rare and pediatric disease clinical trials.
method Dynamic borrowing method based on frequentist approach using similarity measures.
result Demonstrates usefulness of dynamic borrowing in reanalyzing clinical trial data.

Crowdsourced investigation shows differing results for technical analysis strategies.

problem Differing results in technical analysis strategies due to lack of method.
method Collaborative scientific computational framework using Monte Carlo simulations and historical back testing.
result Results are not repeatable by other researchers, highlighting the need for transparency and robustness.

Paper uses AI to predict stock market volatility with neural networks and genetic algorithms.

problem Traditional methods for predicting stock market volatility have high errors.
method Back-propagation neural network and genetic algorithm integrated model.
result The model predicts future volatility with low errors and high accuracy.

Forward translation improves neural machine translation for sentences originally in source language.

problem Improving neural machine translation quality using synthetic data.
method Case study with French-English news translation, separating test sets by original language, analyzing domains, translationese, and noise.
result Forward translation delivers superior gains on sentences originally in source language, complementing back-translation on target language sentences.

The study compares different models for predicting factor premiums and finds neural networks perform better but have unstable weights.

problem Predicting and timing the CMA factor premium using machine learning models.
method Compared regression models (OLS, Ridge, Random Forest, Neural Network) and tested factor timing strategies.
result Neural networks outperform linear models in explaining factor premium variance, but weights are unstable.

Deep learning fails in classifying handwritten historical documents, traditional methods perform better.

problem Classifying handwritten historical documents using deep learning methods.
method Traditional and deep learning methods were tested on a large collection of handwritten historical manuscripts.
result Deep learning methods performed poorly, traditional methods performed consistently well.

In the past 20 years, momentum or trend following strategies have become an established part of the investor toolbox. We introduce a new way of analyzing momentum strategies by looking at the information ratio (IR, average return divided by standard deviation). We calculate the theoretical IR of a momentum strategy, an…

2014-02-13abs ↗pdf ↗

Study optimizes portfolio allocation policies using off-policy data and constraints.

problem Optimizing portfolio allocation policies under constraints using off-policy data.
method Solves a minimax objective with off-policy estimators and online learning to control constraint violations.
result Constructs near-optimal allocation policies for various regimes of operation and constraints.

Study assesses drought and late-frost risks in Bavaria using vine copulas.

problem Assessing risks of late-frost and drought in Bavaria due to climate change.
method Used vine copula models for non-Gaussian and asymmetric dependencies, with univariate and bivariate regression analyses.
result Identified 'at-risk' regions for forest adaptation.

Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy that can significantly reduce the number of labeled samples needed to train a font…

2016-01-27abs ↗pdf ↗

Novel pairs trading strategy for cointegrated cryptocurrencies using copulas.

problem Identifying profitable trading opportunities in cointegrated cryptocurrency pairs.
method Linear and non-linear cointegration tests, correlation coefficient, copula families, back-testing.
result The strategy outperforms buy-and-hold trading strategies in profitability and risk-adjusted returns.

Adaptive Stress Testing detects financial fraud by simulating potential failures.

problem Detecting and mitigating vulnerabilities in financial systems.
method Developed a simplified model using historical data and reinforcement learning.
result Identified the most likely path to system failure and improved fraud detection.

Study efficient sequential evaluation of large language models using historical data.

problem Sequentially evaluate a new large language model (LLM) on a fixed question set.
method Construct a confidence sequence (CS) and design active querying rules to shrink CS width.
result Simple uniform sampling can sometimes outperform adaptive querying rules.

We test a historical price time series in a financial market (the NASDAQ 100 index) for a statistical property known as detailed balance. The presence of detailed balance would imply that the market can be modeled by a stochastic process based on a Markov chain, thus leading to equilibrium. In economic terms, a positiv…

2014-03-14abs ↗pdf ↗

Deep RL model optimizes dynamic portfolio allocation.

problem Automating dynamic portfolio optimization with machine learning.
method Model-based deep reinforcement learning with prediction, data augmentation, and behavior cloning modules.
result Robust, profitable, and risk-sensitive trading strategy compared to baselines.

Although deep learning has historical roots going back decades, neither the term "deep learning" nor the approach was popular just over five years ago, when the field was reignited by papers such as Krizhevsky, Sutskever and Hinton's now classic (2012) deep network model of Imagenet. What has the field discovered in th…

2018-01-02abs ↗pdf ↗

This paper reviews and compares deep generative models for financial time series and VaR.

problem Forecasting risk factor distribution in financial markets.
method Apply multiple deep generative models (CGAN, CWGAN, Diffusion, Signature WGAN) and propose new methods for conditional time series generation.
result Top performing models are Historical Simulation, GARCH, and CWGAN.

Combines historical and market data for better portfolio selection.

problem Improving portfolio selection through diverse information integration.
method Bayesian learning via Gaussian mixture model to harmonize historical and market data.
result The method enhances forecasting accuracy and robustness across various capital markets.

Study uses neural networks to predict stock prices and tests market efficiency.

problem Predicting stock prices from historical data.
method Used Recurrent Neural Networks and Multilayer Perceptrons, compared normalization techniques.
result Found that neural networks can predict stock prices accurately and challenged the efficient-market hypothesis.

Contextualizing financial news improves stock price predictions.

problem Predicting stock prices from financial news requires understanding historical context.
method Proposed a method using a large language model for main articles and a small model for historical context.
result Historical context significantly improves model performance across methods and time horizons.

The paper develops a new framework for detecting distributional drifts conditioned on context.

problem Detecting distributional drifts in machine learning systems when context changes.
method Develops a framework using two-sample tests for conditional distributional treatment effects.
result Demonstrates effectiveness for detecting drift in subpopulations of data.

This study evaluates a dynamic pairs trading strategy in cryptocurrencies using cointegration tests.

problem Improving profitability and risk management in cryptocurrency trading.
method Engle-Granger, KSS, Johansen tests; optimal look-back window; mean-reversion speed calibration; microstructure limitations consideration.
result The strategy outperforms naive buy-and-hold in Bitmex exchange with low maximum drawdown.

GANs can learn stylized facts of financial time series, but performance varies by architecture.

problem Capturing stylized facts of financial time series using GANs.
method Examination of GANs' ability to learn stylized facts of financial time series, focusing on univariate and multivariate data.
result GANs can capture stylized facts of financial time series, but performance varies by architecture.

`Biologically inspired' activation functions, such as the logistic sigmoid, have been instrumental in the historical advancement of machine learning. However in the field of deep learning, they have been largely displaced by rectified linear units (ReLU) or similar functions, such as its exponential linear unit (ELU) v…

2018-03-19abs ↗pdf ↗

This study examines whether tokenized assets improve liquidity and finds significant differences across categories.

problem Improving liquidity for real-world assets through tokenization.
method Examined tokenized real-world assets using Ethereum-based data, measuring liquidity through turnover, active addresses, and active-month indicator.
result Gold-backed tokens show more persistent on-chain activity than Treasury and private-credit-related products, but asset value alone does not reliably predict liquidity.

The statistical analysis of discrete data has been the subject of extensive statistical research dating back to the work of Pearson. In this survey we review some recently developed methods for testing hypotheses about high-dimensional multinomials. Traditional tests like the χ2χ^2 test and the likelihood ratio test ca…

2017-12-17abs ↗pdf ↗

Study evaluates LLMs for predicting Chinese stock movements using financial news sentiments.

problem Evaluating LLMs' ability to predict stock price movements using financial news sentiments.
method Standardized experimental procedure with three LLMs, each with unique performance enhancement methods.
result Developed quantitative trading strategies and conducted back-tests to assess LLMs' performance.

Study finds mean reversion strategies perform well on historical data but fail in recent market conditions.

problem Performance of mean reversion strategies in recent market data.
method Empirical investigation of three mean reversion strategies (PAMR, OLMAR, TCO) on historical S&P 500 data and benchmark datasets.
result Mean reversion strategies may fail in recent market conditions, especially with transaction costs.

The paper uses machine learning to predict missing yield parameters from liquid markets to illiquid corporate bonds.

problem Predicting missing yield parameters from illiquid corporate bonds.
method Applying Denoising Autoencoder (DAE) algorithm to historical data of liquid market instruments.
result DAE algorithm outperforms point-in-time inpainting algorithms in predicting unobserved yield surfaces.

We present an arbitrage-free non-parametric yield curve prediction model which takes the full (discretized) yield curve as state variable. We believe that absence of arbitrage is an important model feature in case of highly correlated data, as it is the case for interest rates. Furthermore, the model structure allows t…

2012-03-09abs ↗pdf ↗

Generative Adversarial Networks simulate realistic market interactions.

problem Lack of agent-level historical data limits market simulation realism.
method Conditional Generative Adversarial Networks (CGANs) trained on real data.
result CGAN-based synthetic market generator outperforms previous methods in market responsiveness and realism.

New method uses Diffusion Maps for latent space modeling of dynamical systems.

problem Building reduced dynamical models from time series data.
method Two rounds of Diffusion Maps on latent coordinates, with lifting back to ambient space.
result Approximation of full state functions in reduced coordinates.

A trading system uses LLMs to adapt to volatile crypto markets.

problem Volatility and market sentiment in cryptocurrencies make traditional models ineffective.
method Specialized LLM agents for technical analysis, sentiment evaluation, and decision-making; verbal feedback for continuous improvement.
result Agents outperform buy-and-hold strategy with consistent gains across market phases.

This study compares Bitcoin and Litecoin using cryptocurrency metrics and trading strategies.

problem Valuation and trading strategies for cryptocurrencies.
method Metrics like UTXO, STXO, WAL, CDD, and trading strategies based on PU ratio.
result Bitcoin's superior store-of-value proposition compared to Litecoin validated.

Artificial neural networks are most commonly trained with the back-propagation algorithm, where the gradient for learning is provided by back-propagating the error, layer by layer, from the output layer to the hidden layers. A recently discovered method called feedback-alignment shows that the weights used for propagat…

2016-09-06abs ↗pdf ↗