Defines interpretability in machine learning and calls for rigorous evaluation.
problem Lack of consensus on what interpretability is and how to measure it.
method Defines interpretability, suggests a taxonomy for rigorous evaluation, and identifies open questions.
result Calls for a more rigorous science of interpretable machine learning.
The authors advocate for more rigorous unsupervised cross-lingual learning methods.
problem Lack of parallel data for many languages.
method Review of unsupervised cross-lingual learning approaches and methodological issues.
result A scenario without parallel data and abundant monolingual data is unrealistic.
A spin network is a cubic ribbon graph labeled by representations of SU(2). Spin networks are important in various areas of Mathematics (3-dimensional Quantum Topology), Physics (Angular Momentum, Classical and Quantum Gravity) and Chemistry (Atomic Spectroscopy). The evaluation of a spin network is an integ…
Synthetic experiments are crucial for assessing causal machine learning methods.
problem Current empirical evaluations of causal machine learning methods are insufficient and unreliable.
method Propose principles for conducting rigorous empirical analyses with synthetic data.
result Rigorous synthetic experiments are essential for building trust in causal machine learning methods.
New method for rigorous confidence intervals in off-policy evaluation.
problem Evaluate new policies from off-policy data without executing them.
method Variational framework using kernel Bellman loss.
result Efficient method for tight confidence intervals in various settings.
SMEs provide a transparent testbed for RL evaluation.
problem Lack of precise, white-box diagnostics in RL environments.
method Synthetic Monitoring Environments (SMEs) with fully configurable task characteristics and known optimal policies.
result SMEs allow for precise evaluation of RL algorithms, revealing the impact of specific environmental properties.
Backward exploration reduces sample complexity in policy evaluation.
problem Empirical policy evaluation in reinforcement learning.
method Backward exploration algorithms from high-cost states.
result Reduced average-case sample complexity to O(logS). New rigorous uncertainty bounds for Gaussian Process regression.
problem Need for frequentist uncertainty bounds in applications like learning-based control.
method Introduce new uncertainty bounds that are rigorous and practically useful.
result New bounds are less conservative and more useful for practical applications.
Stochastic models analyze traffic network performance.
problem Evaluate traffic system performance.
method Stochastic cell transmission models, preference functionals, Gaussian process regression.
result Illustrated in two case studies.
Time series are ubiquitous, and a measure to assess their similarity is a core part of many computational systems. In particular, the similarity measure is the most essential ingredient of time series clustering and classification systems. Because of this importance, countless approaches to estimate time series similar…
This paper evaluates LLMs for technical market analysis, finding GPT-4 Turbo and FinGPT outperform passive benchmarks.
problem Evaluating LLMs for technical market analysis in financial markets.
method Structured evaluation of five LLMs (GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, FinGPT) on four tasks: candlestick pattern recognition, directional signal generation, backtesting, and financial report comprehension.
result GPT-4 Turbo and FinGPT outperform passive benchmarks in simulated backtesting, with GPT-4 Turbo achieving the highest annualized return and Sharpe ratio.
We provide faster algorithms for the problem of Gaussian summation, which occurs in many machine learning methods. We develop two new extensions - an O(Dp) Taylor expansion for the Gaussian kernel with rigorous error bounds and a new error control scheme integrating any arbitrary approximation method - within the best …
The paper evaluates and benchmarks electricity price forecasting models.
problem Lack of rigorous evaluation methods and open datasets.
method Literature review, cross-market comparison, open datasets, and python toolbox.
result Best practices for electricity price forecasting are proposed.
The paper constructs semistrict monoidal 2-categories from foam evaluations.
problem Creating examples of semistrict monoidal 2-categories.
method Using a closed foam evaluation formula as input, the paper rigorously constructs semistrict monoidal 2-categories.
result The constructed monoidal 2-categories are semistrict, have duals and adjoints, and carry a spatial duality structure.
Paper benchmarks adversarial robustness methods on image classification.
problem Vulnerability of deep neural networks to adversarial examples.
method Established a comprehensive benchmark with robustness curves.
result Found important findings on adversarial attack and defense methods.
Noise titration benchmarks time series forecasting models rigorously.
problem Evaluation of time series forecasting models is often flawed due to lack of interventionist methods.
method Interventionist benchmarking using Gaussian noise titration of dynamical systems.
result Fern model outperforms state-of-the-art models in non-stationary conditions.
A rigorous ML pipeline for binary classification in biomedical studies, focusing on pancreatic cancer.
problem Handling bias in ML models for complex biomedical data.
method Customizable ML analysis pipeline with 9 algorithms, hyperparameter optimization, and thorough evaluation.
result Comparison of ML algorithms to ExSTraCS, highlighting interpretability and bias handling.
Paper proposes a sequential statistical test for comparing imitation learning policies with near-optimal stopping.
problem Challenges in rigorously comparing imitation learning policies due to small sample sizes and potential p-hacking.
method Sequential statistical test that adapts the number of trials based on intermediate results, achieving near-optimal stopping.
result Reduces the number of evaluation trials by up to 32% compared to state-of-the-art baselines, saving significant time and effort.
Machine learning suffers from poor design, data, and evaluation practices.
problem Ad hoc design, poor data hygiene, and lack of statistical rigor in model evaluation.
method Examines the entire machine learning process from design to evaluation, highlighting common pitfalls and providing recommendations.
result Common pitfalls in machine learning research and development are identified and actionable recommendations are provided.
This paper evaluates targeted data poisoning attacks by focusing on the hardest samples, improving evaluation and defense strategies.
problem The effectiveness of targeted data poisoning attacks is often overestimated due to average evaluation methods.
method The paper introduces metrics to identify the hardest and easiest to poison samples based on clean model information.
result The proposed metrics reliably stratify samples by poisoning vulnerability, enabling rigorous worst-case evaluation and proactive defense.
Authors provide a fair comparison of GNNs for graph classification.
problem Lack of reproducibility and rigorousness in experimental procedures for GNNs.
method Controlled and uniform framework with over 47,000 experiments.
result GNNs do not fully exploit structural information on some datasets.
Survey on RL reproducibility using real-world robots.
problem Difficulty in reproducing RL results due to algorithm variance, environment stochasticity, and hyper-parameters.
method Investigates issues leading to irreproducible RL research and standardizes evaluation approach.
result Shows how to manage and standardize evaluation of RL algorithms for unbiased comparison.
Develops a validated trading framework for market microstructure signals.
problem Overfitting and lookahead bias in algorithmic trading.
method Interpretable hypothesis-driven signal generation, reinforcement learning, strict out-of-sample testing.
result Modest annualized returns with strong downside protection and market-neutral characteristics.
A hybrid impurity measure balances theoretical soundness and computational efficiency.
problem Developing a robust impurity measure for decision trees.
method Integrates Tsallis entropy with an exponential polarization component.
result Simple parametric measures outperform ITC, but ITC variants are competitive with strong theoretical guarantees.
Quantum method improves CVaR evaluation under correlated fields.
problem Accurately evaluating CVaR in high-dimensional, correlated material uncertainty.
method Quantum-enhanced inference framework using stabilized IQAE.
result Quantum method achieves lower oracle complexity than classical methods.
ERICA assesses reproducibility in cluster analysis.
problem Lack of a unified framework for evaluating cluster analysis replicability.
method ERICA (iterative clustering assignments) method to quantify replicability.
result Demonstrates ERICA's ability to identify reproducible cluster structure.
Graphical lasso models ASR utterance dependencies for consistent WER estimation.
problem Modeling dependent structure among ASR utterances for accurate significance analysis.
method Graphical lasso for dependency modeling, followed by blockwise bootstrap resampling.
result Statistically consistent variance estimator of WER under mild conditions.
Unified approach for non-stationary and clustered bandits.
problem Solving non-stationary and clustered bandits with overlapping solutions.
method Test of homogeneity for seamless integration of non-stationary and clustered bandits.
result Unified solution framework for change detection and cluster identification.
Study evaluates deep learning methods for dermatology, finding they perform poorly under non-ideal conditions.
problem Lack of robustness of deep learning methods in dermatology under real-world conditions.
method Simulated non-ideal conditions on user-submitted dermatology images.
result Deep learning methods show significant drop in accuracy and prediction changes under non-ideal conditions.
Fairmetrics evaluates fairness in ML models for specific groups.
problem Ensuring models do not produce biased outcomes for specific groups.
method User-friendly R package for evaluating group-based fairness criteria.
result Rigorous evaluation of multiple fairness metrics.
FINCH dataset enables financial Text-to-SQL tasks, improving model evaluation.
problem Lack of a large-scale financial dataset for Text-to-SQL tasks.
method Curated financial dataset (FINCH), benchmarking reasoning and language models.
result Proposed FINCH Score for more accurate financial model evaluation.
Automatically evaluates image quality based on human judgment.
problem Difficulty in rigorously evaluating generated image quality.
method Generative model embeddings, human labels regression, and statistical matching.
result 66% accuracy in predicting human scores of image realism.
Interpretable framework evaluates structure learning methods for causal discovery from observational data.
problem Evaluation of structure learning methods under assumption violations in causal discovery.
method Six-dimensional evaluation metric (DOS) tailored for causal discovery.
result Amortized causal discovery delivers results with high proximity to the optimal solution.
Framework for fair predictive models using resampled sensitive attributes.
problem Achieving fair predictions in machine learning models.
method Introducing a discrepancy functional and resampling sensitive attributes.
result Improved performance and equitable uncertainty quantification.
This article evaluates explanations without ground truth in IML.
problem Lack of ground truth for evaluating explanations in IML.
method Rigorously defined problem, reviewed existing efforts, summarized three aspects of explanation, designed a unified evaluation framework.
result Unified evaluation framework for different scenarios in practice.
New measures quantify diversity of latent representations using metric space magnitude.
problem Evaluating the diversity of latent representations in machine learning models.
method Developed magnitude-based measures for latent representations, stable under data perturbations.
result Demonstrated superior performance across various domains and tasks.
The paper develops a new model to evaluate policies in complex temporal/spatial experiments.
problem Evaluating the impact of policies in experiments with temporal and spatial dependencies.
method Temporal/spatio-temporal Varying Coefficient Decision Process (VCDP) model, decomposing ATE into DE and IE.
result Effective estimation and inference of DE and IE with rigorous statistical analysis.
OASIS optimizes ER evaluation by reducing labelling needs with optimal sampling.
problem Extreme class imbalance in ER leads to high labelling costs.
method OASIS uses a biased instrumental distribution and Bayesian updates to focus on unlabelled items.
result OASIS estimates F-measure, precision, recall converge to true values with significant labelling reductions.
Study evaluates LLMs like ChatGPT and GPT-4 on financial analysis tasks.
problem Assessing financial reasoning capabilities of large language models.
method Mock CFA exam questions, Zero-Shot, Chain-of-Thought, Few-Shot scenarios.
result LLMs perform well on financial analysis tasks, but have limitations.
This research improves calorimeter simulations by creating a faster model.
problem Efficiently simulating detailed calorimeter data for high-energy physics.
method Developed a conditional normalizing flow model for superresolution.
result The model successfully reproduces reference distributions.
Develops a method to simulate rare dangerous events in autonomous systems.
problem Rare dangerous events in safety-critical systems are hard to test in real-world settings.
method Combines exploration, exploitation, and optimization techniques for rare-event simulation.
result Provides rigorous guarantees for the performance of the method.
The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.
problem Inaccurate uncertainty estimates in LLM evaluations with small datasets.
method Alternative frequentist and Bayesian methods for uncertainty quantification.
result CLT-based methods underestimate uncertainty in small data settings.
New algorithm maintains privacy while improving model performance in selective release.
problem Privacy degradation and slow convergence in DPSGD.
method Differentially Private Selective Release based on Clipped Gradients (DPSR-CG).
result Maintains strict privacy guarantees while achieving exceptional model performance.
This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…
Adversarial method finds rare catastrophic failures in safety-critical agents.
problem Evaluating safety-critical learning systems for catastrophic failures.
method Adversarial evaluation approach focusing on rare adversarial situations.
result Adversarial evaluation finds catastrophic failures and estimates failure rates faster.
Benchmark assesses LLMs' causal inference skills, revealing significant limitations.
problem Lack of rigorous evaluation of LLMs' causal inference capabilities.
method CausalPitfalls benchmark with structured challenges and grading rubrics.
result Significant limitations in current LLMs' statistical causal inference.
Proposes a method for kernel learning using feature maps.
problem Improving SVM margin through iterative refinement.
method Fourier-analytic characterization and iterative feature maps.
result Optimal and generalization guarantees for SVM margin improvement.
Proposes a method to choose thresholds for LLM evaluation metrics.
problem Ensuring reliable large language models (LLMs) with correct threshold selection.
method Identify risks, stakeholders' risk tolerance, and use ground-truth data to determine thresholds.
result Demonstrates a concrete example with the Faithfulness metric and HaluBench dataset.