Proposes standards for evaluating online machine learning methods in evolving data streams.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study improves off-policy evaluation from non-i.i.d. bandit samples.
Efficiently evaluate generative models at the prompt level using tensor factorization.
We consider evaluation methods for payoffs with an inherent financial risk as encountered for instance for portfolios held by pension funds and insurance companies. Pricing such payoffs in a way consistent to market prices typically involves combining actuarial techniques with methods from mathematical finance. We prop…
We present and implement two algorithms for analytic asymptotic evaluation of the marginal likelihood of data given a Bayesian network with hidden nodes. As shown by previous work, this evaluation is particularly hard for latent Bayesian network models, namely networks that include hidden variables, where asymptotic ap…
FEET protocol evaluates foundation models across three scenarios.
Look-Ahead-Bench evaluates financial LLMs for lookahead bias, revealing significant differences in model performance.
This paper addresses the problem of evaluating learning systems in safety critical domains such as autonomous driving, where failures can have catastrophic consequences. We focus on two problems: searching for scenarios when learned agents fail and assessing their probability of failure. The standard method for agent e…
Evaluation metrics for prediction models don't fully reflect intervention impact.
Automates standards design through machine learning.
Proposes an alternative method for quantifying uncertainty in complex models.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
We present a comparative study of different probabilistic forecasting techniques on the task of predicting the electrical load of secondary substations and cabinets located in a low voltage distribution grid, as well as their aggregated power profile. The methods are evaluated using standard KPIs for deterministic and …
New framework evaluates MIA without retraining, addressing biases.
This research proposes a new (old) metric for evaluating goodness of fit in topic models, the coefficient of determination, or . Within the context of topic modeling, has the same interpretation that it does when used in a broader class of statistical models. Reporting with topic models addresses two c…
Researchers use DL and XAI to evaluate climate downscaling models.
A new method improves AI fairness assessment by estimating performance across intersectional subgroups.
TUDataset provides benchmark datasets for graph learning.
A classical spin network consists of a ribbon graph (i.e., an abstract graph with a cyclic ordering of the vertices around each edge) and an admissible coloring of its edges by natural numbers. The standard evaluation of a spin network is an integer number. In a previous paper, we proved an existence theorem for the as…
New method improves consistency of reinforcement learning performance evaluations.
RobustBench aims to standardize adversarial robustness evaluation in image classification.
Neural network pruning lacks standardized benchmarks and metrics.
We evaluate the hedging performance of a high-order compact finite difference scheme from [4] for option pricing in Bates model. We compare the scheme's hedging performance to standard finite difference methods in different examples. We observe that the new scheme outperforms a standard, second-order central finite dif…
New method stabilizes FQE by reweighting Bellman targets.
New IRT method identifies useful datasets for ML classifier evaluation.
PerturBench benchmarks ML models for cellular perturbation analysis.
Graph neural networks (GNNs) have emerged recently as a powerful architecture for learning node and graph representations. Standard GNNs have the same expressive power as the Weisfeiler-Leman test of graph isomorphism in terms of distinguishing non-isomorphic graphs. However, it was recently shown that this test cannot…
WRENCH benchmarks weak supervision datasets for machine learning.
HED Score improves temporal evaluation of detection accuracy.
We propose a novel adaptive importance sampling algorithm which incorporates Stein variational gradient decent algorithm (SVGD) with importance sampling (IS). Our algorithm leverages the nonparametric transforms in SVGD to iteratively decrease the KL divergence between our importance proposal and the target distributio…
Standard Gaussian Process outperforms in high-dimensional Bayesian Optimization.
Algorithmic risk assessments are increasingly used to help humans make decisions in high-stakes settings, such as medicine, criminal justice and education. In each of these cases, the purpose of the risk assessment tool is to inform actions, such as medical treatments or release conditions, often with the aim of reduci…
FAKI improves gradient-free inference for inverse problems.
State representation learning aims at learning compact representations from raw observations in robotics and control applications. Approaches used for this objective are auto-encoders, learning forward models, inverse dynamics or learning using generic priors on the state characteristics. However, the diversity in appl…
Machine Learning (ML) and Deep Learning (DL) innovations are being introduced at such a rapid pace that model owners and evaluators are hard-pressed analyzing and studying them. This is exacerbated by the complicated procedures for evaluation. The lack of standard systems and efficient techniques for specifying and pro…
ISAAC audits deep models for drug-target interactions, revealing structural differences.
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups…
K-fold cross-validation (CV) with squared error loss is widely used for evaluating predictive models, especially when strong distributional assumptions cannot be taken. However, CV with squared error loss is not free from distributional assumptions, in particular in cases involving non-i.i.d. data. This paper analyzes …
The study improves VaR forecast accuracy by modeling conditional quantile dynamics.
Machine Learning (ML) is increasingly applied in real-life scenarios, raising concerns about bias in automatic decision making. We focus on bias as a notion of opinion exclusion, that stems from the direct application of traditional ML pipelines to infer subjective properties. We argue that such ML systems should be ev…
We consider the asymptotics of the Turaev-Viro and the Reshetikhin-Turaev invariants of a hyperbolic -manifold, evaluated at the root of unity instead of the standard . We present evidence that, as tends to , these invariants grow exponentially with growt…
Evaluating prediction models under covariate shift and selective labels
The paper evaluates and improves uncertainty estimates in neural networks for safety-critical applications.
The paper shows how shared random seeds can reduce variance in machine learning evaluations.
New method reduces diffusion model function evaluations for discrete data.
Machine Learning (ML) and Deep Learning (DL) innovations are being introduced at such a rapid pace that researchers are hard-pressed to analyze and study them. The complicated procedures for evaluating innovations, along with the lack of standard and efficient ways of specifying and provisioning ML/DL evaluation, is a …
Real-world applications often involve domain-specific and task-based performance objectives that are not captured by the standard machine learning losses, but are critical for decision making. A key challenge for direct integration of more meaningful domain and task-based evaluation criteria into an end-to-end gradient…
A new first-order sampler improves diffusion probabilistic model sampling quality.