Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3917821,1731,564 · Jun 202019922001200920172026
← all fields·16 papers on machine learning in Statistical ML · 1 month

Study compares different scoring rules for machine-learned weather forecasts, finding scale-awareness improves forecast realism.

problem Improving the accuracy of machine-learned probabilistic weather forecasts.
method Comparison of scoring rules (CRPS, fair global energy score, graph energy score) and analysis of their impact on forecast field spectra.
result Scale-awareness improves forecast realism, particularly in the tropics.

Paper provides conditions for reliable use of pre-trained embeddings in econometrics.

problem Uncertainty in using pre-trained embeddings for econometric tasks.
method Derives sufficient conditions and convergence rates for machine learning models with pre-trained embeddings.
result Establishes theoretical foundations for reliable use of pre-trained embeddings in econometrics.

Develops a machine learning model to predict ALS progression and assistive device use.

problem Challenges in predicting clinically meaningful milestones in ALS.
method Integrates longitudinal ALSFRS-R trajectories with survival modeling.
result Generates individualized survival curves and predicts wheelchair-free survival.

Microdata improves inflation forecasts after major shocks, study finds.

problem Forecasting inflation in a non-stationary environment with microeconomic data.
method Developed a scan test to detect periods of micro forecast outperformance, combined with adaptive machine learning.
result Micro forecasts improve inflation predictions after major shocks, especially after 2020.

Paper proposes new methods for improving interatomic potentials.

problem Limitations of conventional SO(2) Linear architectures in MLIPs.
method Direct Cartesian construction, recursive Clebsch-Gordan construction, Edge Complex Product Basis, Radial Rotary Complex Attention.
result TECE-OAM-RRA-1.0 achieves SOTA performance on Matbench Discovery.

Spectral deconfounding improves machine learning models by reducing hidden confounding effects.

problem Machine learning models can be misled by hidden confounders, leading to unreliable predictions.
method Develops a nonlinear spectral deconfounding framework for gradient boosting that modifies boosting dynamics to slow down in confounding-aligned directions.
result Spectrally deconfounded boosting improves estimation of the target function under hidden confounding and is more scalable.

A new imputation method estimates missing values by matching observed marginals from masked data.

problem Missing values in data undermine statistical and machine learning analysis.
method Estimates a distribution from masked observations using positive semi-definite kernel density estimation.
result The method yields both single and multiple imputations from the same fitted density, with statistical consistency and fast adaptive excess risk.

Methodology to measure lag relevance in time series models.

problem Measuring lag relevance in machine learning models for univariate time series.
method Ghost variables, Shapley values, additive importance measures, auto-relevance and partial auto-relevance functions, one-step forecast.
result Calculated relevance measures successfully demonstrate expected lag structure in almost all cases.

Optimizes data splitting for shorter conformal prediction intervals.

problem Minimizing prediction interval length while maintaining coverage.
method Theoretical framework for optimal data splitting in split conformal prediction.
result Analytical characterizations of length-optimal split ratios in various settings.

Improves numerical solution of ill-conditioned linear systems for machine learning.

problem Wastefulness and instability in solving ill-conditioned linear systems.
method autonugget combines Richardson extrapolation to determine the solution of the ill-conditioned system, improving accuracy over a single nugget.
result Improves accuracy of numerical solution of ill-conditioned linear systems.

Study identifies and estimates treatment effect heterogeneity within principal stratification subpopulations.

problem Causal inference with intermediate outcomes and treatment effect heterogeneity.
method Proposes a novel doubly cross-fit doubly robust machine learner to efficiently learn conditional principal causal effects under principal ignorability.
result Demonstrates informative patterns of treatment effect heterogeneity within the always-survivor subpopulation in an acute lung injury trial.

Improves full conformal prediction for stochastic non-conformity measures.

problem Inability of existing conditions to guarantee full conformal prediction validity under stochastic settings.
method Introduces a new sufficient condition: Conditional Independence & Permutation Invariance in Distribution.
result Corrects the insufficient condition and provides a new sufficient condition for full conformal prediction validity.

The paper introduces a framework to select efficient datasets for preserving model rankings.

problem Efficient evaluation of machine learning models on small, representative datasets.
method Bootstrap aggregation, clustering, design criteria, random baselines, and greedy farthest-first (FAFI).
result Several selection strategies improve rank preservation compared to random subsets, especially in time series classification.

Machine learning models outperform traditional econometric methods for forecasting term structure of government bonds

problem Forecasting the term structure of government bonds
method Combining traditional econometric models with neural network architectures
result Neural network models consistently outperform traditional models in both forecasting accuracy and portfolio performance

Algorithmic fairness is a field of study that addresses the systematic disadvantage of marginalized groups in machine learning systems.

problem Modern machine learning systems increasingly determine access to economic and social opportunities, leading to structural inequalities and prejudices.
method Statistical and structural approaches to algorithmic fairness.
result The field of algorithmic fairness emerged to address the systematic disadvantage of marginalized groups in machine learning systems.