Backtesting ES and VaR is possible using Diebold-Mariano tests.
problem Eliciting and testing risk measures like ES and VaR.
method Use Diebold-Mariano tests for ES, based on Fissler and Ziegel (2015).
result ES can be jointly elicitable with VaR.
Foundation models improve volatility forecasting in finance.
problem Improving volatility forecasting in financial markets.
method Evaluation of TimesFM model, incremental fine-tuning, comparison with econometric benchmarks.
result Incremental fine-tuning improves forecast accuracy and outperforms traditional models.
New method allows backtesting of systemic risk forecasts.
problem Systemic risk measures are not elitable and identifiable, making backtesting impossible.
method Introduces multi-objective elicitability and Diebold--Mariano type tests.
result Proposes a traffic-light approach for backtesting.
Model predicts short-term Amazon rainforest fires with high accuracy.
problem Accurate short-term forecasting of Amazon rainforest fires is challenging.
method Used Seasonal and Trend decomposition based on Loess combined with multi-month-ahead load forecasting algorithms.
result Proposed decomposition-ensemble models provide more accurate forecasts than other models.
The study improves load forecasting for electricity consumers using advanced machine learning models.
problem Improving short-term load forecasting for effective scheduling and decision-making.
method Proposes and evaluates statistical nonlinear models, including LSTM and GRU, for 15-min frequency electricity load forecasting.
result Advanced models outperform other models in out-of-sample forecasting accuracy, as shown by the Diebold-Mariano test.
Predictions are issued on the basis of certain information. If the forecasting mechanisms are correctly specified, a larger amount of available information should lead to better forecasts. For point forecasts, we show how the effect of increasing the information set can be quantified by using strictly consistent scorin…
A study shows that a fine-tuned model's directional accuracy in financial forecasting is largely due to chance, not skill.
problem Misleading directional accuracy in financial forecasting models.
method A reproducible, frozen-data benchmark with paired significance tests to separate skill from base-rate artifact.
result Fine-tuned models do not show significant directional skill over a base rate of 70% in financial forecasting.
Paper predicts Indian stocks using news psycholinguistic features.
problem Predicting Indian stock market performance using financial news.
method Hybrid intelligent models using psycholinguistic variables (LIWC and TAALES) from news articles.
result GMDH and GRNN are statistically the best techniques for prediction.
Bitcoin price prediction models fail to outperform a simple 'today's price' baseline, especially at longer horizons.
problem Lack of robust models that consistently outperform a naive price predictor at various horizons.
method Surveyed peer-reviewed papers, categorized by evaluation methodology, contrasted with social media discourse, and proposed methodological standards.
result No peer-reviewed study has shown robust superiority over the naive baseline across multiple market regimes at short-to-medium horizons.
Study evaluates multivariate forecasting scoring rules and proposes new copula-based ones.
problem Evaluating and improving multivariate probabilistic forecasting methods.
method Analysis and comparison of existing scoring rules, development of copula-based scoring rules, simulation studies, and real data analysis.
result Proposed copula-based scoring rules provide strong distinction between models with correct and incorrect dependency structures.
New method evaluates AI stock prediction systems based on decision-making processes.
problem Lack of evaluation for AI systems' decision-making processes.
method Scores intermediate decision process using large language models and closed-loop reinforcement learning feedback.
result Composite behavioral score correlates with Sharpe ratio and reduces prediction error.
Pretrained time-series models outperform train-from-scratch baselines in financial return forecasting.
problem Financial return forecasting
method Pretrained time-series foundation models
result Pretrained TSFMs dominate the ranking distribution, accounting for 8 of 10 task-level wins.
THRML uses energy-based models for index tracking, reducing portfolio tracking error and improving returns.
problem NP-hard combinatorial optimization in portfolio optimization under cardinality constraints.
method THRML reformulates index tracking as probabilistic inference on an Ising Hamiltonian, using GPU-accelerated block Gibbs sampling.
result THRML achieves 4.31 percent annualized tracking error compared to 5.66-6.30 percent for baselines, with 128.63 percent total return.
Our model predicts stock market intervals using chaotic fusion and graph convolutional networks.
problem Uncertainty in financial market predictions without quantified uncertainty.
method Bi-level chaotic fusion, graph convolutional networks, volatility-aware gating, temporal dependencies.
result Significant improvements in prediction intervals and coverage compared to existing methods.
Thinking LLMs struggle with stock prediction, especially as data complexity increases.
problem Evaluating the performance of 'thinking' LLMs in stock prediction, especially under varying levels of cross-sectional complexity.
method Rolling 48m/1m walk-forward evaluation, comparing direct LLMs, TLLMs, and classical learners on cross-sectional ranking loss, MSE, and backtests with transaction costs.
result TLLMs' ranking quality deteriorates as cross-sectional complexity grows, while direct LLMs remain stable.
Proposes provenance and pseudo-provenance for automated test generation.
problem Invalidation of provenance in generated tests.
method Annotation of generated tests with provenance trails and pseudo-provenance.
result Validates the reliability of generated tests and their relation to seeds.
Survey of Machine Learning Testing: Properties, Components, and Trends.
problem Challenges in testing machine learning models.
method Comprehensive review of 144 ML testing papers.
result Identification of research challenges and directions.
A family of maximum mean discrepancy (MMD) kernel two-sample tests is introduced. Members of the test family are called Block-tests or B-tests, since the test statistic is an average over MMDs computed on subsets of the samples. The choice of block size allows control over the tradeoff between test power and computatio…
USP test improves on Pearson's chi-squared and G-test for independence.
problem Deficiencies in Pearson's chi-squared and G-test for independence. method USP test based on U-statistic estimator of population dependence measure. result USP test controls size, handles small cell counts, and detects minimal violations of independence.
E-C2ST uses E-values for high-dimensional data two-sample tests.
problem Statistical testing for high-dimensional data.
method Combines split likelihood ratio tests and predictive independence tests, using E-values for anytime-valid sequential tests.
result E-C2ST achieves enhanced statistical power by partitioning datasets into multiple batches.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
problem Efficiently testing distribution differences and independence.
method Group datapoints into bins and permute only these bins, using stored sufficient statistics.
result Cheap permutation tests maintain the accuracy and optimality of standard tests but are significantly faster.
Paper proposes a chi-square test for distance correlation.
problem Testing distance correlation is computationally expensive.
method Proposes a chi-square test for distance correlation, non-parametric, fast, applicable to various metrics.
result Chi-square test exhibits similar power to permutation test and can be valid and universally consistent for testing independence.
The paper tests properties of multiple distributions with limited samples.
problem Testing properties of multiple distributions with few samples.
method Designing testers for uniformity, identity, and closeness testing under specific conditions.
result Sample optimal testers for uniformity, identity, and closeness testing are provided.
Optimizes two-sample tests for non-Euclidean domains using spectral regularization.
problem Optimizing two-sample tests for non-Euclidean domains.
method Spectral regularization of MMD test to achieve minimax optimality.
result Proposes a spectral regularization method that improves test optimality.
DRIFT uses RL to automate functional software testing efficiently.
problem Efficient and reliable automated software testing.
method DRIFT employs Q-learning with Graph Neural Networks on symbolic UI representations.
result DRIFT can robustly test software functionalities in a fully automated manner.
Detects overfitting in models trained on test sets.
problem Challenges in verifying overfitting without independent test sets.
method Uses adversarial examples and unbiased error estimates to test for overfitting.
result Correctly identifies overfitting to the training set but not to the test set.
Post hoc test for Sharpe ratio improves pairwise comparisons.
problem Improving pairwise comparisons of Sharpe ratios.
method Analogous to Tukey's test, applied after rejecting equal Signal-Noise ratios.
result Maintains nominal type I rate and is moderately powerful.
Robust test for distributions under Hellinger distance, simpler than optimal tests.
problem Testing and estimating distributions robustly under Hellinger distance.
method Simple robust hypothesis test with optimal sample complexity, robust to Hellinger distance perturbations.
result Empirically demonstrated robustness and power of the test on canonical distributions.
Framework for online hypothesis testing across various data types.
problem Testing various nonparametric hypotheses in data streams.
method Unified framework using operators on data distributions, leveraging ML models.
result Efficient, adaptive, and error-controlled sequential tests.
A new method for kernel tests without data splitting increases power.
problem Lack of power in kernel-based tests due to data splitting.
method Selective inference framework to learn hyperparameters and test on full sample.
result Empirically larger test power without data splitting, regardless of split proportion.
Develops a new test for comparing two groups' densities, showing minimax optimality.
problem Comparing probability densities between two groups.
method Probabilistic tensor product smoothing spline framework for joint density modeling; penalized likelihood ratio test for interaction testing.
result Proposed test is minimax optimal and outperforms conventional approaches.
Model-X test detects conditional independence in streaming data.
problem Detecting conditional independence in data streams with arbitrary dependency.
method Sequential testing inspired by model-X and testing by betting.
result Significantly reduces type-I error rate and enhances data efficiency.
Simple methods combine statistical tests for out-of-distribution detection.
problem Detecting data points not following the training distribution.
method Combining classical parametric tests (Rao's score test) and a typicality test.
result Combining Fisher's method of test statistics improves out-of-distribution detection accuracy.
Deep neural networks improve two-sample testing.
problem Efficiently distinguishing between two unknown distributions.
method Deep learning representations for two-sample testing.
result Significant reduction in type-2 error rate compared to existing methods.
Unified score and distance-based GoF tests for model adequacy.
problem Difficulty in extending score-based GoF tests to nonparametric alternatives.
method Introducing semiparametric kernelized Stein discrepancy (SKSD) test.
result SKSD test is computationally efficient and universally consistent.
Two new tests approximate KCIT for fast CI testing in large datasets.
problem CI testing is slow and resource-intensive for large datasets.
method Developed RCIT and RCoT, which approximate KCIT using random Fourier features.
result RCIT and RCoT scale linearly with sample size and return accurate p-values faster than KCIT.
The study evaluates various jump tests for high-frequency financial data.
problem Choosing the most effective jump test for high-frequency financial data.
method An extensive evaluation of multiple alternative tests in various scenarios.
result Guidelines for choosing the most suitable test for different settings.
New methods for CI testing under model misspecification.
problem Challenges in CI testing with misspecified models.
method Proposes new approximations and upper bounds for testing errors of regression-based CI tests.
result Introduces the Rao-Blackwellized Predictor Test (RBPT) robust against misspecified inductive biases.
AutoML simplifies two-sample tests for detecting distribution shifts.
problem Detecting distribution shifts between datasets.
method Uses mean discrepancy of a witness function with squared loss minimization.
result AutoML simplifies and improves two-sample testing performance.
New KCM tests improve specification testing via RKHS.
problem Improving specification tests for econometric models.
method Kernel conditional moment (KCM) tests based on RKHS.
result KCM tests have better finite-sample performance than existing tests.
New graph tests improve on existing methods for comparing large graphs.
problem Comparing large graphs from different sources.
method Proposed new tests based on asymptotic distributions.
result New tests are computationally less expensive and more reliable.
Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
problem Comparing Multiscale Fisher's Independence Test (MultiFIT) to HSIC tests for multivariate dependence.
method Compares MultiFIT to HSIC tests, highlighting exact level control and performance limitations.
result Observes performance limitations of MultiFIT in terms of test power.
Develops hypothesis tests for conditional distributions using learning-theoretic bounds.
problem Testing differences in conditional distributions and functionals.
method Transforming learning-theoretic bounds into hypothesis tests for conditional expectations.
result Establishes comprehensive foundation for conditional testing, including theoretical guarantees and practical implementations.
The paper improves the robustness of approximate randomization tests.
problem Noisy data limits the robustness of approximate randomization tests.
method Derives non-asymptotic bounds and novel conditions for approximate randomization tests.
result Valid approximate randomization tests under data invariances can be derived.
Paper explores limits of graph property testing in Ising models.
problem Understanding information-theoretic limits of graph property testing in Ising models.
method Combinatorial constructs and correlation-based tests.
result Property testing is more challenging for general Ising models than ferromagnets.
Develops tests for conditional symmetry under group actions.
problem Testing conditional symmetry in distributions under group actions.
method Nonparametric randomization tests with kernel methods and asymptotic consistency.
result Tests achieve finite-sample Type I error control and power.
DeepEvolution improves testing of deep neural networks by generating diverse test cases.
problem Lack of diversity in generated test cases from random fuzzing or transformations.
method Search-based approach using metaheuristics to ensure diversity in test cases.
result DeepEvolution significantly increases neuronal coverage and detects latent defects.
Two-sample tests using analytic functions are faster and more powerful than previous methods.
problem Testing differences between two probability distributions efficiently.
method Tests based on analytic functions representing distributions, using either smoothed empirical characteristic functions or embeddings in a reproducing kernel Hilbert space.
result The new tests are faster and more powerful than previous methods, detecting differences almost surely at finite locations/frequencies.