Study constructs a Japanese financial LLM benchmark.
problem Need for domain-specific benchmarks for LLMs.
method Constructed a benchmark with multiple Japanese and financial domain tasks.
result GPT-4 outperforms other models in the benchmark.
DLBricks automates DL benchmarking on CPUs, reducing effort and time.
problem Lack of representative and up-to-date DL benchmarks on CPUs.
method Decomposes DL models into runnable networks, leveraging layer repetition and auto-generation.
result Accurately estimates DL model performance and speeds up benchmarking time.
Develops a new method for benchmark portfolios and market outperformance strategies.
problem Creating effective benchmark portfolios for market outperformance.
method Explicit formulaic algorithm and multifactor risk model tailored for long-only portfolios.
result Explicit positive weights for benchmarks without principal components or iterations.
Procgen Benchmark uses procedurally generated games to test reinforcement learning.
problem Lack of diverse and high-quality training environments for reinforcement learning.
method Developed 16 procedurally generated game-like environments and used them to benchmark reinforcement learning.
result Procedurally generated environments are essential for training and evaluating reinforcement learning agents.
Benchmark for DL inference on embedded HWAs, focusing on autonomous driving.
problem Lack of comprehensive benchmarks for DL hardware.
method Developed a benchmark for inference on embedded HWAs, focusing on autonomous driving. Proposed new granularity, benchmark procedures, and performance indicators.
result Identifies mismatches between HWAs and DL models.
Fidel-TS creates a new benchmark for time series forecasting models.
problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.
Tiny benchmarks reduce LLM evaluation costs by using fewer examples.
problem Expensive evaluation of LLMs with tens of thousands of examples.
method Developed evaluation tools and tiny versions of popular benchmarks.
result Accurately estimate LLM performance with just 100 curated examples.
The study initiates a theoretical analysis of dynamic benchmarking models.
problem Lack of theoretical foundation and empirical studies in dynamic benchmarks.
method Examined two realizations of dynamic benchmarking: sequential and hierarchical dependency models.
result Sequential dynamic benchmarks show initial performance improvement but can stall after three rounds due to label noise.
A new sparse benchmark metabench identifies key abilities from large benchmarks.
problem Redundancy and compression in existing benchmarks.
method Data from 5000+ LLMs to identify most informative items, distilling a sparse benchmark.
result Sparse benchmark metabench captures underlying abilities with high accuracy.
MLPerf benchmark suite evaluates diverse ML applications, highlighting system bottlenecks.
problem Understanding and optimizing ML applications across different models and systems.
method Analysis of MLPerf benchmark suite characteristics and comparison with previous benchmarks.
result Optimal distributed deep learning training requires dedicated low latency interconnects.
New meta-score EPP interprets model performance differences.
problem Lack of interpretable benchmarks for model performance.
method Elo-based Predictive Power (EPP) meta-score, logistic regression.
result EPP scores have probabilistic interpretation and can be compared between data sets.
Public benchmark for machine learning models in critical care.
problem Lack of public benchmarks for machine learning in critical care.
method Defined four tasks (mortality prediction, length of stay, phenotyping, decompensation risk) and compared clinical and deep learning models on eICU dataset.
result First public benchmark on multi-centre critical care dataset, comparing clinical models with predictive models.
Surrogate benchmarks improve AC evaluation efficiency.
problem Hard and computationally expensive AC benchmarks.
method Construct surrogate benchmarks from AC benchmarks using regression models.
result Surrogate benchmarks capture AC scenario characteristics efficiently.
Benchmarking deep time series models for equity portfolios
problem Selecting the best deep time series model for equity portfolios
method Using a CRSP benchmark and multi-criteria acceptability analysis
result No architecture dominates the benchmark, with TransEnc-8 having the highest rank-1 acceptability
A scalable DL benchmarking platform for evaluating and comparing models, frameworks, and hardware.
problem Lack of a uniform DL benchmarking platform for evaluating and comparing innovations.
method Identified 10 design features for a DL benchmarking platform, proposed MLModelScope, and implemented as an open-source project.
result Demonstrated how model, hardware, and framework selection affect accuracy and performance under different scenarios.
Language model benchmarks often misrepresent true understanding, revealing vulnerabilities in evaluation methods.
problem Language model benchmarks fail to accurately reflect true language understanding and adaptability.
method Systematic analysis of NLP evaluation frameworks, identifying vulnerabilities in static benchmarks, human evaluation protocols, and LLM-as-judge frameworks.
result Current evaluation methods are unreliable and need improvement to accurately assess LLM performance.
New benchmarks for RNA 3D structure-function modeling.
problem Lack of standardized benchmarks for RNA deep learning.
method Developed seven benchmark datasets, provided tools for data handling, and offered a user-friendly environment for model comparison.
result Demonstrated utility with baseline results using a relational graph neural network.
We solve the multi-criteria benchmarking problem by formalizing it as a social choice problem and identifying conditions for meaningful rankings.
problem Aggregating multiple metrics into a single ranking for models in benchmarking problems.
method Formalizing multi-criteria benchmarking as a social choice problem and identifying sufficient conditions for meaningful rankings.
result We prove that meaningful multi-criteria benchmarking becomes possible under certain preference conditions (single-peaked, group-separable, distance-restricted).
Benchmarking TPU, GPU, and CPU for deep learning models.
problem Improving performance in deep learning training.
method ParaDnn benchmark suite for FC, CNN, and RNN models on TPU, GPU, and CPU.
result TPU, GPU, and CPU have unique strengths for different types of models.
GIFT-Eval benchmarks time series forecasting models across diverse datasets.
problem Lack of comprehensive benchmarks for evaluating time series foundation models.
method Developed GIFT-Eval, a benchmark with 23 datasets, 177 million data points, and 144,000 time series.
result Promotes evaluation of foundation models across various domains and frequencies.
Benchmark proposes to assess molecule docking efficiency.
problem Lack of realistic benchmarks for measuring progress in drug design.
method Proposes a docking-based benchmark using SMINA software.
result Graph-based generative models fail to generate high-scoring molecules.
BREEDS benchmarks assess model robustness to subpopulation shifts.
problem Measuring model robustness to novel subpopulation shifts.
method Controlled synthesis of realistic distribution shifts using class structure.
result Validated model sensitivity and effectiveness of robustness interventions.
YAHPO Gym introduces a new benchmark for evaluating hyperparameter optimization methods.
problem Evaluating and comparing hyperparameter optimization methods on well-curated benchmark suites.
method Surrogate-based benchmark collection of 14 scenarios, each with multi-fidelity and multi-objective hyperparameter optimization problems.
result Surrogate-based benchmarks produce more faithful results than tabular benchmarks.
New benchmark for non-rigid 3D human shape retrieval.
problem Distinguishing between body shapes of 3D human models.
method Extended benchmark with 145 new models and FAUST dataset.
result Improved comparison of 25 shape retrieval methods.
Benanza speeds up DL model optimization by automatically generating micro-benchmarks and identifying inefficiencies.
problem Slow characterization/optimization cycles for DL models on GPUs.
method Benanza includes a model processor, benchmark generator, database, and analyzer.
result Benanza identifies optimizations in parallel layer execution, cuDNN, framework inefficiency, layer fusion, and Tensor Cores.
Paper proposes a new speech representation benchmark and model.
problem Lack of benchmarks for comparing speech representations.
method Unsupervised triplet-loss objective for training a universal non-semantic speech representation.
result Proposed representation outperforms other models on benchmark and transfer learning tasks.
We argue for the principle of unchanged optimality in RL benchmarks and discuss its implications.
problem Generalization in reinforcement learning benchmarks.
method Discussion of conceptual properties and subtle choices in state representation and model architecture.
result The principle of unchanged optimality is important for RL benchmarks and can be broken or satisfied by model architecture choices.
We study the pricing and hedging of derivatives in incomplete financial markets by considering the local risk-minimization method in the context of the benchmark approach, which will be called benchmarked local risk-minimization. We show that the proposed benchmarked local risk-minimization allows to handle under extre…
Benchmark tests spoken language models for infant language learning.
problem Understanding how infants learn language from speech.
method Developed a language-acquisition-friendly benchmark.
result Benchmarking shows models' strengths and weaknesses.
OceanForecastBench offers a comprehensive benchmark for data-driven ocean forecasting models.
problem Lack of open-source, standardized benchmarks for data-driven ocean forecasting models.
method Proposes OceanForecastBench, a benchmark with high-quality data and evaluation pipeline.
result Offers the most comprehensive benchmarking framework for data-driven ocean forecasting.
Benchmark improves object detection robustness in winter weather.
problem Assessing object detection models' performance under image corruptions.
method Developed three benchmark datasets with various image corruptions; used data augmentation to improve robustness.
result Simple data augmentation significantly enhances model robustness across different corruptions and datasets.
PerturBench benchmarks ML models for cellular perturbation analysis.
problem Standardizing benchmarking in modeling single cell transcriptomic responses to perturbations.
method Modular platform, diverse datasets, metrics, extensive evaluation, rank metrics.
result Simpler models are competitive and scale well with larger datasets.
UCFE benchmarks LLMs in financial tasks with human feedback.
problem Evaluating LLMs' financial task performance and user satisfaction.
method Hybrid approach combining human expert evaluations and dynamic interactions.
result Significant alignment between benchmark scores and human preferences (Pearson correlation coefficient of 0.78).
Study benchmarks 19 survival models on 34 datasets, finding Cox model still best.
problem Quantitative comparison of survival models on low-dimensional data.
method Comprehensive benchmarking of 19 models on 34 datasets, tuning and evaluating using 6 metrics.
result Cox Proportional Hazards model remains best overall for low-dimensional, right-censored data.
Paper benchmarks MBRL algorithms with new environments.
problem Lack of standardized benchmarking for MBRL algorithms.
method Proposes 18 new benchmarking environments and standardizes problem settings.
result Characterizes three key research challenges for MBRL.
Bayesian framework for comparing trading algorithms using cost analysis.
problem Comparing trading algorithms using cost analysis.
method Bayesian framework, hierarchical models, standardized benchmarks, fat tails, skewness, heteroscedasticity.
result Effective calculation of trading benchmarks with limited data.
RealCause provides a realistic benchmark for causal inference.
problem Lack of a reliable benchmark for comparing causal effect estimators.
method Flexible generative models to create a benchmark that is both ground-truth and realistic.
result Evaluation of over 1500 causal estimators provides evidence for choosing hyperparameters using predictive metrics.
Benchmark assesses forecasting models' ability to use textual context.
problem Forecasting models struggle with integrating textual context.
method Introduces a benchmark with numerical and textual data, evaluates various models.
result LLM prompting method outperforms other models.
We developed a framework to benchmark and compare competing risks survival models.
problem Limited systematic evaluation and adoption of competing risks survival models.
method Open-source reproducible benchmarking framework for comprehensive comparison.
result Systematic comparison across multiple datasets on various performance aspects.
SD-SCMs generate counterfactual data for causal inference benchmarks.
problem Benchmarking causal inference methods with realistic data.
method Sequence-driven structural causal models (SD-SCMs) for causal inference.
result State-of-the-art methods struggle with individual treatment effect estimation.
Paper introduces benchmark-neutral pricing for long-term contracts.
problem High prices of long-term contracts under risk-neutral pricing.
method Uses growth optimal portfolio as numeraire and new pricing measure.
result Identifies minimal possible prices for contingent claims.
New benchmark for earthquake forecasting models shows current neural point processes are not yet suitable.
problem Lack of a modern benchmark for evaluating neural point process models in earthquake forecasting.
method Curated and standardized earthquake catalog, evaluation protocols, and datasets.
result None of the tested NPPs outperformed the classical ETAS model.
LOB-Bench benchmarks generative AI for financial data, outperforming traditional models.
problem Lack of consensus on evaluating generative AI models for financial data.
method Python-based benchmark with LOB statistics and market impact metrics.
result Generative autoregressive models outperform traditional models in LOB data.
The paper clarifies conditions for using benchmark scores in machine learning.
problem Using benchmark scores to draw scientific inferences about learning problems.
method Developing conditions of construct validity inspired by psychological measurement theory.
result Clarifies conditions under which benchmark scores support diverse scientific claims.
Pythae is a Python library for benchmarking VAE models.
problem Improving variational autoencoders for various tasks.
method Unified implementation and framework for 19 generative autoencoder models.
result Benchmarking 19 VAE models across multiple tasks.
Benchmark for math reasoning models from human proofs.
problem Measuring and accelerating machine learning models in high-level mathematical reasoning.
method Built a non-synthetic dataset from theorem prover proofs, defined a task for model to fill in missing propositions, used hierarchical transformer to improve performance.
result Neural models can capture non-trivial mathematical reasoning, hierarchical transformer outperforms baseline.
The paper shows that benchmark-neutral pricing minimizes option prices.
problem Pricing extreme-maturity European put options on diversified indices.
method Benchmark-neutral pricing applied to a drifted time-transformed squared Bessel process.
result Benchmark-neutral price is the minimal possible price, risk-neutral price is more expensive.
Sloth predicts LLM performance using latent skills across families.
problem Variations in benchmark performance due to differences in training configurations and data processing across model families.
method Sloth uses publicly available benchmark data and assumes LLM performance is driven by latent skills influenced by model size and training tokens. It exploits correlations across benchmarks to provide accurate predictions.
result Sloth predicts LLM performance accurately and offers insights into scaling behaviors for complex tasks.