Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3917821,1731,564 · Jun 202019922001200920172026
← all fields·60 papers on machine learning in Quant Finance · 1 year

Quantum kernels show no advantage in stock return prediction, but differ in stability metrics.

problem Determining if quantum kernels improve stock return prediction.
method Controlled horse race on Chinese A-share market with identical training subsamples and tuning budgets.
result Quantum kernels do not outperform classical RBF controls in cross-sectional stock return prediction.

AI models failed to profitably predict cryptocurrency extrema on Binance Spot.

problem Tackling the profitability of candle-based machine learning models for short-term cryptocurrency trading.
method Scripted fixed-seed model runs and deterministic simulators with human supervision.
result Strongest evidence found negative, with models underperforming buy-and-hold strategies.

OMD monitors stock market dynamics through matrix trajectories and reveals crisis patterns.

problem Understanding and predicting stock market crises and sector rotations.
method Applying OMD to S\&P 500 returns over three crises, analyzing distance matrices and their spectra.
result Market dynamics show coherent changes during crises, with distinct sector leadership.

SHARC explains machine learning risk models for regulatory capital, linking outputs to scenarios.

problem Inability to explain machine learning model outputs to regulatory bodies.
method SHAP-based explainability framework for Hybrid GPR-HS architecture and SVaR stress-testing.
result SHARC links SVaR outputs to scenario inputs, providing auditable traceability.

Machine learning models outperform traditional econometric methods for forecasting term structure of government bonds

problem Forecasting the term structure of government bonds
method Combining traditional econometric models with neural network architectures
result Neural network models consistently outperform traditional models in both forecasting accuracy and portfolio performance

Hierarchical graph learning for calendar spread strategies in commodity futures markets

problem Developing machine-learning methods for calendar spread strategies in commodity futures markets
method Proposing a hierarchical graph learning approach
result Outperforming benchmark models in both prediction and trading performance

A Longitudinal Attribute-Conditioned Neural Network (LANTERN) framework for modeling health-state transition probabilities in irregular longitudinal data.

problem Estimating long-term care transition probabilities in irregular longitudinal health data.
method A neural network that learns from individual health history, incorporates time elapsed, and conditions on demographic and socioeconomic attributes.
result Improves severe disability discrimination and maintains strong calibration.

Research shows ESG signals lower exposure to market fragility during stress periods.

problem Market fragility often occurs together, and ESG is associated with reduced exposure.
method Monthly data on S&P 500 constituents from 2014 to 2025, analyzing downside returns, volatility, illiquidity, and cofragility states.
result A one-standard-deviation increase in ESG lowers the probability of severe cofragility by 0.92 percentage points during stress periods.

Machine learning predicts Bitcoin returns but trading performance drops with costs.

problem Trading Bitcoin predictions with transaction costs.
method XGBoost, LSTM, iTransformer models evaluated in walk-forward protocol; cost-aware execution filter implemented.
result Cost-aware execution filter restores profitability; XGBoost strategy outperforms.

Breaks circular dependency in synthetic option pricing with a novel model.

problem Circular dependency in implied volatility limits synthetic data for machine learning and risk analysis.
method Uses a Jump-Hidden Markov Model to generate price paths and a modified Heston process to convert paths into implied volatility.
result Framework generates realistic synthetic American option prices without external calibration.

Transformer pre-training improves stock return prediction accuracy.

problem Improving stock price prediction accuracy for better investment decisions.
method Pre-trained transformer models on TSX index, fine-tuned for individual stocks, compared to LSTM and XGBoost.
result Transformer model achieved lower mean squared error than benchmarks.

Benchmark detects decision-time leakage in financial backtests.

problem Detecting decision-time leakage in financial machine-learning backtests.
method Toggles one evaluation convention at a time around a clean t+1t{+}1-open reference, holding other factors fixed.
result Inflation is highly selective, affecting specific features and execution methods.

The paper develops a new framework for managing asymmetric volatility.

problem Managing asymmetric volatility to improve recovery and participation.
method Path-dependent framework for asymmetric volatility management.
result Skew engineering reduces harmful downside participation more than productive upside participation.

Hybrid method improves SABR implied volatility approximation.

problem Improving SABR implied volatility approximation.
method Combining analytical structure with machine learning, using geometric features and residual correction.
result Hybrid model improves accuracy and robustness compared to analytical and neural-network approaches.

PHBench predicts Series A funding from Product Hunt launch signals with 7.8% accuracy.

problem Predicting startup Series A funding from launch signals on Product Hunt.
method Constructed PHBench from 67,292 Product Hunt posts, linked to funding records, and used a three-component ensemble model.
result Best-performing model achieved F0.5 = 0.097 and AP = 0.037, with a statistically significant advantage over logistic regression.

Machine learning improves beta forecasts, enhancing equity valuation and portfolio performance.

problem Improving beta forecasts for better equity valuation and portfolio performance.
method Using machine learning on a large cross-section of US stocks with various firm characteristics.
result Machine learning improves out-of-sample performance of asymmetric beta measures.

Study improves motor insurance claim prediction using geographic data.

problem Limited location identifiers in public actuarial datasets.
method Zone-level modeling framework with environmental and orthoimagery data.
result Geographic information improves MTPL claim prediction accuracy.

Quantum computing offers new solutions for financial optimization, pricing, risk, and security.

problem Core financial bottlenecks in combinatorial search, expectation estimation, and rare-event analysis.
method Identify bottlenecks, specify quantum primitives, compare with classical benchmarks, assess under constraints.
result Strongest near-term case for quantum finance in hybrid workflows, constrained search, and amplitude-estimation.

The paper proposes a machine learning framework for portfolio optimization with limited data.

problem Low data environments and regime uncertainty in portfolio optimization.
method A teacher-student learning pipeline with CVaR optimizer generating supervisory labels and neural models trained on real and synthetic data.
result Student models can match or outperform the CVaR teacher and achieve improved robustness under regime shifts.

A method for accurate pricing of multidimensional derivatives under uncertain volatility.

problem High-dimensional stochastic control problem in uncertain volatility model.
method Backward actor-critic stochastic policy gradient scheme combining DP, PPO, and neural networks.
result Accurate and efficient pricing of multidimensional derivatives compared to benchmarks.

Adapts Altman's model to compositional data for bankruptcy prediction.

problem Predicting business default using standard financial ratios has issues.
method Uses compositional data methodology with log-ratios and machine learning.
result Compositional methods improve predictive performance, especially random forests.

Paper uses bipartite graph to forecast cross-market returns, revealing asymmetry.

problem Cross-market return predictability and asymmetry between U.S. and Chinese markets.
method Directed bipartite graph capturing time-ordered linkages, hypothesis testing for edge selection, regularized and ensemble machine learning models.
result U.S. returns predict Chinese intraday returns, but not vice versa, revealing asymmetry.

Improved deep hedging with ensemble uncertainty quantification.

problem Uncertainty in deep hedging models hinders their deployment.
method Trained an ensemble of LSTM networks to quantify uncertainty in deep hedging under Heston volatility and proportional transaction costs.
result The ensemble's disagreement provides a strong predictive confidence measure for hedge performance.

A machine learning method for short-maturity options with jumps and stochastic volatility.

problem Short-maturity options with jumps and stochastic volatility.
method Differential machine learning method combining supervision and PIDE-residual penalty.
result Improves jump-term approximation and reduces Greeks errors compared to baselines.

The paper shows how overreactions in stock prices can be predicted and used for trading.

problem Predicting and monetizing overreactions in stock prices as momentum signals.
method High-frequency data from Twitter, machine learning models (XGBoost, Random Forests, Deep Neural Networks, Bidirectional LSTMs), and SHAP for explainability.
result Machine learning models significantly outperform traditional overreaction rules at ultra short horizons.

Factor Engine simplifies financial factor computation and analysis in Python.

problem Efficient computation and analysis of financial factors.
method Modular, extensible Python library with decorators, integrates with data science ecosystem.
result Mispricing factors computed by Factor Engine and Stata implementation are highly similar.

BPASGM uses sparse graphical models to optimize portfolio selection.

problem Portfolio optimization in high-dimensional settings with estimation error.
method BPASGM extends BPA to a sparse graphical model, screening assets for diversification.
result BPASGM portfolios outperform standard mean-variance portfolios in risk-adjusted performance.

Generative AI improves stock selection by synthesizing features from diverse data sources.

problem Automating feature discovery in stock market data.
method Used large language models with retrieval-augmented generation and structured prompting to synthesize features from various data sources.
result AI-generated features consistently outperform baselines, with Sharpe improvements ranging from 14% to 91%.

Hybrid AI system combines technical, sentiment analysis for adaptive equity trading.

problem Traditional trading strategies fail during high volatility and regime shifts.
method Combines trend-following, mean-reversion, sentiment analysis, machine learning, and market regime filtering.
result Hybrid model achieved 135.49% return on investment over 24 months.

Paper compares econometric models with machine learning for energy forecasting.

problem Tackles the trade-off between predictive accuracy and interpretability in energy markets.
method Integrates TVP-SVAR with copulas for forecasting energy--macro dynamics.
result Copula-enhanced econometric models provide interpretable insights while matching machine learning accuracy.

Develops a machine-learning framework for optimal share repurchase hedging.

problem Challenges in hedging share repurchase programs due to market regulations and trading activity.
method Machine-learning framework that optimizes execution and hedging of share repurchase programs.
result Substantial performance improvements and an optimized hedging approach.

GT-Score reduces overfitting in trading strategies by integrating multiple criteria.

problem Overfitting in data-driven financial models leads to unreliable out-of-sample performance.
method Integrates performance, statistical significance, consistency, and downside risk into a composite objective function.
result Improves generalization ratio by 98% compared to baseline objective functions in walk-forward validation.

Paper explores balancing market dynamics and interpretable forecasting models for energy prices.

problem Tackles the challenge of accurately predicting mFRR price and understanding market dynamics.
method Compares XGBoost and EBM for forecasting mFRR activation price in the balancing market.
result EBM provides comparable forecasting accuracy to XGBoost but with higher interpretability.

Fossil power firms have recently profited more than renewables, but this may be a temporary phenomenon.

problem The profitability gap between renewable and fossil power firms in Europe.
method Machine-learning clustering and Bayesian model averaging.
result Renewable power firms are becoming more profitable, while fossil power firms are becoming less so.

XGBoost predicts NEPSE Index log returns with low error and high directional accuracy.

problem Forecasting daily log-returns in the NEPSE Index with high accuracy.
method XGBoost machine learning, feature engineering, hyperparameter optimization, walk-forward validation.
result Optimal XGBoost configuration achieves lowest log-return RMSE and MAE.

New method cleans cross-covariance matrices for better financial forecasting.

problem Asymptotically optimal cross-covariance cleaners fail in real-world, time-varying markets.
method Physics-informed neural network that learns from empirical singular values.
result Trained model outperforms analytical cleaners in out-of-sample cross-covariance prediction.

A robust machine learning approach forecasts U.S. Treasury yields, reducing risk for investors.

problem Noisy and uncertain U.S. Treasury yields pose risk to forecast users.
method Formulates yield curve forecasting as a distributionally robust problem, combining factor models and machine learning.
result Robust forecast combinations improve out-of-sample performance across different maturity periods.

Study integrates climate and text data to improve credit default prediction.

problem Improving credit risk assessment for mSEs with limited financial histories.
method Multimodal framework using LSTM, GRU, and transformer models.
result Integration of multiple data modalities improves credit default prediction.

New method improves stock return prediction in non-stationary markets.

problem Tackles the challenge of predicting stock returns in non-stationary environments.
method Jointly optimizes model class and training window size using a tournament procedure.
result Consistently outperforms standard benchmarks by 14-23% in out-of-sample R2R^2.

The study uses machine learning to predict CAT bond coupons based on climate data.

problem Predicting CAT bond coupons using climate data.
method Combining climate indicators with machine learning models (random forest, gradient boosting, etc.).
result Extremely randomized trees achieved the lowest RMSE in predicting CAT bond coupons.

Study predicts bond yields using machine learning and ultimate forward rates.

problem Forecasting bond yields using ultimate forward rates.
method Applied de Kort-Vellekooptype methodology for UFR estimation, used linear and nonlinear machine learning techniques.
result Nonlinear machine learning models outperform linear models in bond yield forecasting.

A hybrid framework uses machine learning to price options faster and more accurately.

problem Rapid recalibration of option pricing models in dynamic markets.
method Integrates smooth offset algorithm with supervised machine learning models.
result Surrogate pricing operators achieve up to 1000x speedup over direct SOA evaluation.

This paper benchmarks monotone-constrained models for credit PD across datasets and finds constraints are mostly costless.

problem Aligning machine learning model behavior with domain knowledge in credit risk.
method Benchmarked monotone-constrained versus unconstrained gradient boosting models across five datasets and three libraries, defining the Price of Monotonicity (PoM) as the relative change in AUC.
result Monotonicity constraints are almost costless on large datasets and most costly on smaller datasets, with PoM ranging from essentially zero to about 2.9 percent.

Study enhances financial forecasting with machine learning and fuzzy MCDM.

problem Increasing financial uncertainty and market complexity.
method Integrates machine learning (XGBoost, LSTM, GNN) and intuitionistic fuzzy MCDM.
result High forecasting accuracy with low MAPE and narrow confidence intervals.

Paper combines LSTM and Random Forest for better stock market predictions.

problem Improving stock market trading predictions by integrating technical and fundamental data.
method Integrates LSTM networks with Random Forest algorithms using financial and microeconomic data.
result Hybrid approach outperforms traditional methods combining both technical and fundamental variables.

Machine learning struggles to predict binary options movements due to randomness.

problem Predicting binary options movements using machine learning.
method Tested multiple machine learning models (RF, LR, GB, kNN) and neural networks (MLP, LSTM) on EUR/USD currency pairs.
result None of the models surpassed the ZeroR baseline accuracy, indicating randomness in binary options.

New framework allows selective removal of stale data in option calibration.

problem Inability to remove old data from calibrated option pricing models without full retraining.
method Introduces operator-theoretic Gauss-Newton framework for selective forgetting.
result Provides stability guarantees and perturbation bounds for selective data removal.

Study assesses impact of CBDC on financial stability in dual-currency economy.

problem Impact of CBDC on financial stability in dual-currency economy (Romania).
method Integrated analytical framework combining econometrics, machine learning, and behavioural modelling. CBDC adoption probabilities estimated using XGBoost and logistic regression models. Liquidity stress simulations and VAR, MSVAR, SVAR models capture macro-financial transmission.
result CBDC uptake would be moderate, primarily driven by digital readiness and trust in the central bank.

A machine learning approach for dynamic stock recommendation outperforms traditional strategies.

problem Lack of time for analysts to check all S&P 500 stocks and the need for a reliable stock selection strategy.
method Selecting representative stock indicators, using five machine learning methods, and choosing the model with the lowest Mean Square Error to rank stocks.
result The proposed scheme outperforms the long-only strategy on the S&P 500 index in terms of Sharpe ratio and cumulative returns.

Study examines cross-training neural networks for financial index prediction.

problem Predicting financial indexes from different markets using machine learning.
method Investigated various neural network architectures and trained them on one market index to predict another.
result Cross-training models on one market index improved prediction accuracy for another market index.