ModelRadar evaluates forecasting models across multiple aspects.
problem Evaluating forecasting models using single scores hides relevant performance variations.
method ModelRadar, a framework for aspect-based evaluation of univariate time series forecasting models.
result NHITS performs best overall but its superiority varies with forecasting conditions.
New method evaluates language model forecasters by checking consistency of predictions.
problem Evaluating the performance of language model forecasters is difficult due to lack of ground truth.
method Developed a consistency check framework based on arbitrage to evaluate forecasters.
result Consistency metrics correlate with ground truth performance of LLM forecasters.
The paper evaluates forecast accuracy of realized volatility measures in large cross-sections.
problem Forecast evaluation of realized volatility measures in large cross-sections of financial data.
method Equal predictive accuracy testing procedures, LASSO shrinkage, measurement error correction, cross-sectional jump component measures.
result The augmented HAR model outperforms the standard HAR model in forecasting realized volatility.
A new framework evaluates deep learning vs classical forecasting methods for time series predictions.
problem Current forecasting model evaluation metrics fail to capture model performance differences.
method Proposes a novel framework for evaluating univariate time series forecasting models from multiple perspectives.
result Deep learning models like NHITS outperform classical methods in multi-step ahead forecasting but not in anomaly handling.
The article reviews scoring rules for estimating and evaluating forecasts.
problem Evaluating probabilistic forecasts and estimating probability distributions.
method Mathematical foundations and characterization of scoring rules.
result Important families of scoring rules and their applications in statistics and machine learning.
The problem of automatic and accurate forecasting of time-series data has always been an interesting challenge for the machine learning and forecasting community. A majority of the real-world time-series problems have non-stationary characteristics that make the understanding of trend and seasonality difficult. Our int…
This paper evaluates forecast quality in electricity markets beyond traditional accuracy measures.
problem Traditional accuracy measures fail to reflect the economic value of electricity price forecasts.
method Investigates four quality dimensions: accuracy, dispersion, association, and extremum identification.
result Dispersion- and association-based measures better capture forecast economic value.
EasyTime simplifies time series forecasting for researchers and practitioners.
problem Ease of use and accuracy in time series forecasting.
method One-click evaluation, automated ensemble, natural language Q&A.
result Superior forecasting accuracy compared to individual methods.
Study identifies regions where scoring rules reliably detect forecast errors.
problem Insufficient reliability of scoring rules in evaluating multivariate probabilistic forecasts.
method Systematic finite-sample analysis of proper scoring rules on synthetic and real-world data.
result Identified regions of reliability for scoring rules in time-series forecasting.
Archive of 20 time series datasets for forecasting evaluation.
problem Lack of comprehensive time series forecasting datasets.
method Compilation and characterisation of 20 datasets from various domains.
result Characterisation and performance evaluation of datasets.
This work shows how evaluation metrics can be seen as fair gambles.
problem The relationship and evaluation of machine learning forecasts.
method Using game-theoretic probability, the authors show evaluation metrics as fair gambles.
result Standard evaluation metrics are fair gambler outcomes, with calibration and regret metrics on two dimensions.
The study evaluates forecast risk-adjusted performance using various metrics.
problem Evaluating forecast reliability beyond accuracy.
method Risk-adjusted performance measures (Sharpe, Sortino, Omega ratios) and Edge Ratio.
result Machine learning models often offer attractive risk profiles but not necessarily higher reliability.
The primary goal of this study is doing a meta-analysis research on two groups of published studies. First, the ones that focus on the evaluation of the United States Department of Agriculture (USDA) forecasts and second, the ones that evaluate the market reactions to the USDA forecasts. We investigate four questions. …
The study evaluates financial risk using copulas and statistical tests.
problem Validating bivariate forecasts in risk evaluation.
method Using copulas to characterize dependencies, applying statistical tests to validate forecasts, removing heteroskedasticity.
result A Student copula accurately describes financial time series dependencies.
Fidel-TS creates a new benchmark for time series forecasting models.
problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.
Conditional forecasts improve performative prediction accuracy.
problem Performative predictions undermine standard forecasting methods.
method Condition forecasts on covariates to make them forecast-invariant.
result Proper scoring rules fail under conditioning, but two solutions are identified.
OceanForecastBench offers a comprehensive benchmark for data-driven ocean forecasting models.
problem Lack of open-source, standardized benchmarks for data-driven ocean forecasting models.
method Proposes OceanForecastBench, a benchmark with high-quality data and evaluation pipeline.
result Offers the most comprehensive benchmarking framework for data-driven ocean forecasting.
We present a comparative study of different probabilistic forecasting techniques on the task of predicting the electrical load of secondary substations and cabinets located in a low voltage distribution grid, as well as their aggregated power profile. The methods are evaluated using standard KPIs for deterministic and …
FinTSBridge evaluates financial time series models for asset pricing.
problem Lack of effective evaluation methods for financial time series models.
method Developed FinTSBridge suite with new metrics and tasks.
result Showcased new metrics for financial time series models.
This paper challenges the current metrics used for evaluating long-term forecasting models.
problem Current metrics focus on pointwise error reduction, ignoring structural properties.
method Proposes a multi-dimensional evaluation approach that includes statistical fidelity, structural coherence, and decision-level relevance.
result Current progress in forecasting may reflect specialization in benchmark configurations rather than deeper understanding of temporal dynamics.
Kernel quadrature improves CRPS estimation for probabilistic time-series forecasting.
problem Intractable integrations in CRPS evaluation metrics lead to improper rankings of forecasting models.
method Introduced kernel quadrature approach for unbiased CRPS estimation and scalable computation.
result Our approach consistently outperforms existing CRPS estimators.
Develops RES metrics for stable rare-event forecasting evaluation.
problem Challenges in evaluating forecasts of rare events.
method Rare-event-stable (RES) metrics designed to maintain stable thresholds under extreme rarity.
result RES metrics maintain stable thresholds, consistent model rankings, and near-complete prevalence invariance.
Generative model improves intraday electricity price forecasting.
problem Intraday electricity price forecasting for improved trading strategies.
method Generative neural network model for probabilistic path forecasts.
result Generative model leads to higher profit gains than benchmark methods.
The study analyzes how probabilistic forecasts improve battery trading strategies in electricity markets.
problem Improvements in statistical forecast quality do not directly translate to economic value in battery trading strategies.
method The study frames battery optimization as a stochastic program based on fully probabilistic forecasts and examines decision quality under different uncertainty models.
result The study identifies two critical flaws in quantile-based trading strategies and provides theoretical justification and empirical evidence.
Hybrid approach improves probabilistic forecasts for electricity trading.
problem Improving probabilistic forecasts for electricity trading markets.
method Combines QRA and factor-based averaging for probabilistic forecasting.
result The hybrid approach outperforms benchmarks in statistical measures and economic value.
AI agents improve forecast combination in empirical economics.
problem Hidden researcher degrees of freedom in AI-generated code.
method Adapted agent-loop architecture to empirical economics, added holdout evaluation.
result Independent agent searches find better forecast methods than benchmarks.
Study evaluates local explanation methods for time series forecasting.
problem Lack of local interpretability methods for multivariate time series forecasting.
method Proposed two novel evaluation metrics: Area Over the Perturbation Curve for Regression and Ablation Percentage Threshold.
result Comprehensive comparison of local explanation models on two datasets.
The study improves VaR forecast accuracy by modeling conditional quantile dynamics.
problem Improving the accuracy of Value-at-Risk (VaR) forecasts for time-varying quantiles.
method Time-varying modeling of VaR, evaluation via simulation, asymmetric Mean Absolute Deviation loss function.
result Substantial improvements in forecasting conditional quantiles by maintaining predicted quantile unchanged.
An image-driven approach for time series forecasting.
problem Time-series forecasting as a computer vision task.
method Capture input data as an image and train a model to produce the subsequent image.
result Our method outperforms various baselines, including ARIMA, using image-based evaluation metrics.
GIFT-Eval benchmarks time series forecasting models across diverse datasets.
problem Lack of comprehensive benchmarks for evaluating time series foundation models.
method Developed GIFT-Eval, a benchmark with 23 datasets, 177 million data points, and 144,000 time series.
result Promotes evaluation of foundation models across various domains and frequencies.
AI agents improve forecast combination but require transparency.
problem AI coding agents increase flexibility in empirical economics, leading to hidden degrees of freedom.
method Adapted open-source agent-loop architecture to empirical economics workflow, adding post-search holdout evaluation.
result Multiple agent runs outperform standard benchmarks in rolling evaluation but not all on post-search holdout.
Study improves distributional regression evaluation with CRPS, finding optimal rates of convergence.
problem Improving probabilistic forecasts in meteorology using distributional regression.
method Extends theoretical properties of CRPS evaluation to include covariates and finite sample sizes, analyzing convergence rates for different methods.
result Optimal minimax rate of convergence for distributional regression methods is achieved by k-nearest neighbor and kernel methods.
New methods for scoring function decomposition improve forecast evaluation.
problem Improving forecast evaluation and understanding forecast components.
method Linear recalibration of forecasts for miscalibration, discrimination, and uncertainty.
result Enhanced statistical power and deeper insights into forecast components.
Motivated by the Basel 3 regulations, recent studies have considered joint forecasts of Value-at-Risk and Expected Shortfall. A large family of scoring functions can be used to evaluate forecast performance in this context. However, little intuitive or empirical guidance is currently available, which renders the choice…
New metrics improve probabilistic forecasting, especially for rare events.
problem Current evaluation frameworks for probabilistic forecasting assume independence and lack sensitivity to tail events.
method Proposed signature kernel-based metrics: Sig-MMD and CSig-MMD.
result These metrics capture complex dependencies and prioritize tail event prediction.
Unified NICEk metrics improve solar forecasting accuracy.
problem Lack of suitable error metrics for multidimensional solar irradiance forecasting.
method Introducing NICEk framework with Lk norms for evaluating forecasting models.
result NICESigma consistently outperforms traditional metrics in discriminative power and statistical significance.
In order to obtain a reasonable and reliable forecast method for crude oil price volatility, this paper evaluates the forecast performance of single-regime GARCH models (including the standard linear GARCH model and the nonlinear GJR-GARCH and EGARCH models) and the two-regime Markov Regime Switching GARCH (MRS-GARCH) …
YC Bench forecasts startup success in Y Combinator batches with a short-term metric.
problem Difficult forecasting of startup success due to sparse meaningful outcomes and slow evaluation cycles.
method Developed a live benchmark using publicly available traction signals and web visibility metrics.
result Revealed 6 out of 11 top performers at YC Demo Day with a simple proxy for prior brand recognition.
Time series forecasting is difficult. It is difficult even for recurrent neural networks with their inherent ability to learn sequentiality. This article presents a recurrent neural network based time series forecasting framework covering feature engineering, feature importances, point and interval predictions, and for…
This paper analyzes forecasting models for COVID-19 cases and deaths.
problem Reliable forecasting of COVID-19 cases and deaths is crucial for managing the disease.
method Quantitative analysis of forecasting models across different regions in the US, evaluating model selection, hyperparameter tuning, and training time.
result Model selection is the most influential factor in forecasting performance.
Cisco introduces a new time series model for better forecasting.
problem Improving time series forecasting accuracy.
method Developed a new multiresolution decoder-only model trained on large datasets.
result The new model achieves superior performance on observability datasets.
Study evaluates various regularization methods for electricity price forecasting.
problem Improving accuracy of electricity price predictions.
method Applied ten different penalty functions to two model structures in two electricity markets.
result LQ and elastic net consistently produce more accurate forecasts than other regularization types.
Deep neural networks forecast financial return distributions accurately.
problem Forecasting probability distributions of financial returns.
method Used 1D CNN and LSTM architectures with custom loss functions to optimize distribution parameters.
result LSTM with skewed Student's t distribution outperformed classical models in multiple evaluation metrics.
The paper evaluates and benchmarks electricity price forecasting models.
problem Lack of rigorous evaluation methods and open datasets.
method Literature review, cross-market comparison, open datasets, and python toolbox.
result Best practices for electricity price forecasting are proposed.
Research combines econometric, machine learning, and deep learning models for financial forecasting.
problem Improving financial time series forecasting accuracy.
method Hybrid models combining ARIMA, SVM, XGBoost, and LSTM.
result Effective hybrid models outperform individual components and the Buy&Hold strategy.
New benchmark for earthquake forecasting models shows current neural point processes are not yet suitable.
problem Lack of a modern benchmark for evaluating neural point process models in earthquake forecasting.
method Curated and standardized earthquake catalog, evaluation protocols, and datasets.
result None of the tested NPPs outperformed the classical ETAS model.
Chronos models improve financial forecasting by integrating multivariate data.
problem Improving financial forecasting accuracy using multivariate data.
method Evaluation of Chronos-2 on multivariate and univariate financial forecasting models.
result Multivariate forecasts consistently outperform univariate forecasts, especially for interest rates.
Signature kernel scoring rule improves weather forecasting by capturing temporal and spatial dependencies.
problem Lack of suitable scoring rules for probabilistic weather forecasting.
method Reframe weather variables as continuous paths using iterated integrals (signature kernels) to capture temporal and spatial dependencies.
result Signature kernel scoring rule outperforms conventional methods in weather forecasting, especially for long-term forecasts.