Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

106212318424 · Jun 202019922001200920172026
48 results for forecast evaluation

ModelRadar evaluates forecasting models across multiple aspects.

problem Evaluating forecasting models using single scores hides relevant performance variations.
method ModelRadar, a framework for aspect-based evaluation of univariate time series forecasting models.
result NHITS performs best overall but its superiority varies with forecasting conditions.

New method evaluates language model forecasters by checking consistency of predictions.

problem Evaluating the performance of language model forecasters is difficult due to lack of ground truth.
method Developed a consistency check framework based on arbitrage to evaluate forecasters.
result Consistency metrics correlate with ground truth performance of LLM forecasters.

The paper evaluates forecast accuracy of realized volatility measures in large cross-sections.

problem Forecast evaluation of realized volatility measures in large cross-sections of financial data.
method Equal predictive accuracy testing procedures, LASSO shrinkage, measurement error correction, cross-sectional jump component measures.
result The augmented HAR model outperforms the standard HAR model in forecasting realized volatility.

A new framework evaluates deep learning vs classical forecasting methods for time series predictions.

problem Current forecasting model evaluation metrics fail to capture model performance differences.
method Proposes a novel framework for evaluating univariate time series forecasting models from multiple perspectives.
result Deep learning models like NHITS outperform classical methods in multi-step ahead forecasting but not in anomaly handling.

The article reviews scoring rules for estimating and evaluating forecasts.

problem Evaluating probabilistic forecasts and estimating probability distributions.
method Mathematical foundations and characterization of scoring rules.
result Important families of scoring rules and their applications in statistics and machine learning.

This paper evaluates forecast quality in electricity markets beyond traditional accuracy measures.

problem Traditional accuracy measures fail to reflect the economic value of electricity price forecasts.
method Investigates four quality dimensions: accuracy, dispersion, association, and extremum identification.
result Dispersion- and association-based measures better capture forecast economic value.

Study identifies regions where scoring rules reliably detect forecast errors.

problem Insufficient reliability of scoring rules in evaluating multivariate probabilistic forecasts.
method Systematic finite-sample analysis of proper scoring rules on synthetic and real-world data.
result Identified regions of reliability for scoring rules in time-series forecasting.

This work shows how evaluation metrics can be seen as fair gambles.

problem The relationship and evaluation of machine learning forecasts.
method Using game-theoretic probability, the authors show evaluation metrics as fair gambles.
result Standard evaluation metrics are fair gambler outcomes, with calibration and regret metrics on two dimensions.

The primary goal of this study is doing a meta-analysis research on two groups of published studies. First, the ones that focus on the evaluation of the United States Department of Agriculture (USDA) forecasts and second, the ones that evaluate the market reactions to the USDA forecasts. We investigate four questions. …

2018-01-19abs ↗pdf ↗

The study evaluates financial risk using copulas and statistical tests.

problem Validating bivariate forecasts in risk evaluation.
method Using copulas to characterize dependencies, applying statistical tests to validate forecasts, removing heteroskedasticity.
result A Student copula accurately describes financial time series dependencies.

Fidel-TS creates a new benchmark for time series forecasting models.

problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.

Conditional forecasts improve performative prediction accuracy.

problem Performative predictions undermine standard forecasting methods.
method Condition forecasts on covariates to make them forecast-invariant.
result Proper scoring rules fail under conditioning, but two solutions are identified.

OceanForecastBench offers a comprehensive benchmark for data-driven ocean forecasting models.

problem Lack of open-source, standardized benchmarks for data-driven ocean forecasting models.
method Proposes OceanForecastBench, a benchmark with high-quality data and evaluation pipeline.
result Offers the most comprehensive benchmarking framework for data-driven ocean forecasting.

We present a comparative study of different probabilistic forecasting techniques on the task of predicting the electrical load of secondary substations and cabinets located in a low voltage distribution grid, as well as their aggregated power profile. The methods are evaluated using standard KPIs for deterministic and …

2019-10-03abs ↗pdf ↗

This paper challenges the current metrics used for evaluating long-term forecasting models.

problem Current metrics focus on pointwise error reduction, ignoring structural properties.
method Proposes a multi-dimensional evaluation approach that includes statistical fidelity, structural coherence, and decision-level relevance.
result Current progress in forecasting may reflect specialization in benchmark configurations rather than deeper understanding of temporal dynamics.

Kernel quadrature improves CRPS estimation for probabilistic time-series forecasting.

problem Intractable integrations in CRPS evaluation metrics lead to improper rankings of forecasting models.
method Introduced kernel quadrature approach for unbiased CRPS estimation and scalable computation.
result Our approach consistently outperforms existing CRPS estimators.

Develops RES metrics for stable rare-event forecasting evaluation.

problem Challenges in evaluating forecasts of rare events.
method Rare-event-stable (RES) metrics designed to maintain stable thresholds under extreme rarity.
result RES metrics maintain stable thresholds, consistent model rankings, and near-complete prevalence invariance.

Generative model improves intraday electricity price forecasting.

problem Intraday electricity price forecasting for improved trading strategies.
method Generative neural network model for probabilistic path forecasts.
result Generative model leads to higher profit gains than benchmark methods.

The study analyzes how probabilistic forecasts improve battery trading strategies in electricity markets.

problem Improvements in statistical forecast quality do not directly translate to economic value in battery trading strategies.
method The study frames battery optimization as a stochastic program based on fully probabilistic forecasts and examines decision quality under different uncertainty models.
result The study identifies two critical flaws in quantile-based trading strategies and provides theoretical justification and empirical evidence.

Hybrid approach improves probabilistic forecasts for electricity trading.

problem Improving probabilistic forecasts for electricity trading markets.
method Combines QRA and factor-based averaging for probabilistic forecasting.
result The hybrid approach outperforms benchmarks in statistical measures and economic value.

AI agents improve forecast combination in empirical economics.

problem Hidden researcher degrees of freedom in AI-generated code.
method Adapted agent-loop architecture to empirical economics, added holdout evaluation.
result Independent agent searches find better forecast methods than benchmarks.

Study evaluates local explanation methods for time series forecasting.

problem Lack of local interpretability methods for multivariate time series forecasting.
method Proposed two novel evaluation metrics: Area Over the Perturbation Curve for Regression and Ablation Percentage Threshold.
result Comprehensive comparison of local explanation models on two datasets.

The study improves VaR forecast accuracy by modeling conditional quantile dynamics.

problem Improving the accuracy of Value-at-Risk (VaR) forecasts for time-varying quantiles.
method Time-varying modeling of VaR, evaluation via simulation, asymmetric Mean Absolute Deviation loss function.
result Substantial improvements in forecasting conditional quantiles by maintaining predicted quantile unchanged.

GIFT-Eval benchmarks time series forecasting models across diverse datasets.

problem Lack of comprehensive benchmarks for evaluating time series foundation models.
method Developed GIFT-Eval, a benchmark with 23 datasets, 177 million data points, and 144,000 time series.
result Promotes evaluation of foundation models across various domains and frequencies.

AI agents improve forecast combination but require transparency.

problem AI coding agents increase flexibility in empirical economics, leading to hidden degrees of freedom.
method Adapted open-source agent-loop architecture to empirical economics workflow, adding post-search holdout evaluation.
result Multiple agent runs outperform standard benchmarks in rolling evaluation but not all on post-search holdout.

Study improves distributional regression evaluation with CRPS, finding optimal rates of convergence.

problem Improving probabilistic forecasts in meteorology using distributional regression.
method Extends theoretical properties of CRPS evaluation to include covariates and finite sample sizes, analyzing convergence rates for different methods.
result Optimal minimax rate of convergence for distributional regression methods is achieved by k-nearest neighbor and kernel methods.

Motivated by the Basel 3 regulations, recent studies have considered joint forecasts of Value-at-Risk and Expected Shortfall. A large family of scoring functions can be used to evaluate forecast performance in this context. However, little intuitive or empirical guidance is currently available, which renders the choice…

2017-05-12abs ↗pdf ↗

New metrics improve probabilistic forecasting, especially for rare events.

problem Current evaluation frameworks for probabilistic forecasting assume independence and lack sensitivity to tail events.
method Proposed signature kernel-based metrics: Sig-MMD and CSig-MMD.
result These metrics capture complex dependencies and prioritize tail event prediction.

Unified NICEk metrics improve solar forecasting accuracy.

problem Lack of suitable error metrics for multidimensional solar irradiance forecasting.
method Introducing NICEk framework with Lk norms for evaluating forecasting models.
result NICESigma consistently outperforms traditional metrics in discriminative power and statistical significance.

YC Bench forecasts startup success in Y Combinator batches with a short-term metric.

problem Difficult forecasting of startup success due to sparse meaningful outcomes and slow evaluation cycles.
method Developed a live benchmark using publicly available traction signals and web visibility metrics.
result Revealed 6 out of 11 top performers at YC Demo Day with a simple proxy for prior brand recognition.

Time series forecasting is difficult. It is difficult even for recurrent neural networks with their inherent ability to learn sequentiality. This article presents a recurrent neural network based time series forecasting framework covering feature engineering, feature importances, point and interval predictions, and for…

2019-01-01abs ↗pdf ↗

This paper analyzes forecasting models for COVID-19 cases and deaths.

problem Reliable forecasting of COVID-19 cases and deaths is crucial for managing the disease.
method Quantitative analysis of forecasting models across different regions in the US, evaluating model selection, hyperparameter tuning, and training time.
result Model selection is the most influential factor in forecasting performance.

Study evaluates various regularization methods for electricity price forecasting.

problem Improving accuracy of electricity price predictions.
method Applied ten different penalty functions to two model structures in two electricity markets.
result LQ and elastic net consistently produce more accurate forecasts than other regularization types.

Deep neural networks forecast financial return distributions accurately.

problem Forecasting probability distributions of financial returns.
method Used 1D CNN and LSTM architectures with custom loss functions to optimize distribution parameters.
result LSTM with skewed Student's t distribution outperformed classical models in multiple evaluation metrics.

The paper evaluates and benchmarks electricity price forecasting models.

problem Lack of rigorous evaluation methods and open datasets.
method Literature review, cross-market comparison, open datasets, and python toolbox.
result Best practices for electricity price forecasting are proposed.

New benchmark for earthquake forecasting models shows current neural point processes are not yet suitable.

problem Lack of a modern benchmark for evaluating neural point process models in earthquake forecasting.
method Curated and standardized earthquake catalog, evaluation protocols, and datasets.
result None of the tested NPPs outperformed the classical ETAS model.

Chronos models improve financial forecasting by integrating multivariate data.

problem Improving financial forecasting accuracy using multivariate data.
method Evaluation of Chronos-2 on multivariate and univariate financial forecasting models.
result Multivariate forecasts consistently outperform univariate forecasts, especially for interest rates.

Signature kernel scoring rule improves weather forecasting by capturing temporal and spatial dependencies.

problem Lack of suitable scoring rules for probabilistic weather forecasting.
method Reframe weather variables as continuous paths using iterated integrals (signature kernels) to capture temporal and spatial dependencies.
result Signature kernel scoring rule outperforms conventional methods in weather forecasting, especially for long-term forecasts.