Deep learning directly extracts features for emotion recognition, outperforming traditional methods.
problem Emotion recognition using traditional feature encoding methods.
method Used deep learning networks to directly encode features for emotion recognition.
result Highest performance in emotion recognition on EmoDB dataset.
This work tackles domain shift in speech emotion recognition by proposing class-wise adversarial domain adaptation.
problem Domain shift between corpora poses a challenge for speech emotion recognition, especially for positive/negative emotions.
method Class-wise adversarial domain adaptation to reduce shift between different corpora.
result Our method is effective even with limited target labeled examples, as demonstrated on EMODB and Aibo corpora.
Study evaluates feature selection methods for emotion recognition in resource-constrained settings.
problem Reducing memory and computational requirements for emotion recognition in low-resource settings.
method Evaluation of three feature selection methods: ILFS, ReliefF, Fisher, and AFS.
result Smaller feature sets can achieve similar or better accuracy, reducing resource usage.
The study corrects misconceptions in GBDT speed benchmarks.
problem Misleading speed benchmarks of GBDT algorithms.
method Explained and criticized several straightforward benchmarking methods, outlined fair benchmark requirements.
result A fair GBDT speed benchmark requires specific conditions.
YAHPO Gym introduces a new benchmark for evaluating hyperparameter optimization methods.
problem Evaluating and comparing hyperparameter optimization methods on well-curated benchmark suites.
method Surrogate-based benchmark collection of 14 scenarios, each with multi-fidelity and multi-objective hyperparameter optimization problems.
result Surrogate-based benchmarks produce more faithful results than tabular benchmarks.
Study constructs a Japanese financial LLM benchmark.
problem Need for domain-specific benchmarks for LLMs.
method Constructed a benchmark with multiple Japanese and financial domain tasks.
result GPT-4 outperforms other models in the benchmark.
A new sparse benchmark metabench identifies key abilities from large benchmarks.
problem Redundancy and compression in existing benchmarks.
method Data from 5000+ LLMs to identify most informative items, distilling a sparse benchmark.
result Sparse benchmark metabench captures underlying abilities with high accuracy.
Machine learning research depends on objectively interpretable, comparable, and reproducible algorithm benchmarks. We advocate the use of curated, comprehensive suites of machine learning tasks to standardize the setup, execution, and reporting of benchmarks. We enable this through software tools that help to create an…
DLBricks automates DL benchmarking on CPUs, reducing effort and time.
problem Lack of representative and up-to-date DL benchmarks on CPUs.
method Decomposes DL models into runnable networks, leveraging layer repetition and auto-generation.
result Accurately estimates DL model performance and speeds up benchmarking time.
Study optimizes portfolio to minimize relative drawdown duration, penalizing unfavorable performance states.
problem Minimizing relative drawdown duration in portfolio optimization relative to a benchmark.
method Introduces a benchmark-relative drawdown-duration criterion penalizing unfavorable performance states. Uses a one-dimensional Markovian representation and Hamilton-Jacobi-Bellman equation.
result Derives explicit projection-based characterization of the optimal feedback control and identifies geometric settings for unique strong solutions.
Procgen Benchmark uses procedurally generated games to test reinforcement learning.
problem Lack of diverse and high-quality training environments for reinforcement learning.
method Developed 16 procedurally generated game-like environments and used them to benchmark reinforcement learning.
result Procedurally generated environments are essential for training and evaluating reinforcement learning agents.
MLPerf benchmarks ML training to drive performance improvements.
problem Unique challenges in ML training benchmarks.
method Developed MLPerf to overcome ML training's specific challenges.
result Quantitatively evaluated MLPerf's effectiveness.
Deployment-complete benchmarking assesses if evidence leads to consistent deployment actions.
problem Lack of clear evidence leading to consistent deployment actions.
method Introduces deployment-complete benchmarking to test if benchmark evidence determines deployment actions.
result Benchmark evidence must be complete for a claim to lead to a consistent deployment action.
We give an explicit formulaic algorithm and source code for building long-only benchmark portfolios and then using these benchmarks in long-only market outperformance strategies. The benchmarks (or the corresponding betas) do not involve any principal components, nor do they require iterations. Instead, we use a multif…
Tiny benchmarks reduce LLM evaluation costs by using fewer examples.
problem Expensive evaluation of LLMs with tens of thousands of examples.
method Developed evaluation tools and tiny versions of popular benchmarks.
result Accurately estimate LLM performance with just 100 curated examples.
MLPerf benchmark suite evaluates diverse ML applications, highlighting system bottlenecks.
problem Understanding and optimizing ML applications across different models and systems.
method Analysis of MLPerf benchmark suite characteristics and comparison with previous benchmarks.
result Optimal distributed deep learning training requires dedicated low latency interconnects.
The optimization of algorithm (hyper-)parameters is crucial for achieving peak performance across a wide range of domains, ranging from deep neural networks to solvers for hard combinatorial problems. The resulting algorithm configuration (AC) problem has attracted much attention from the machine learning community. Ho…
Benchmark for DL inference on embedded HWAs, focusing on autonomous driving.
problem Lack of comprehensive benchmarks for DL hardware.
method Developed a benchmark for inference on embedded HWAs, focusing on autonomous driving. Proposed new granularity, benchmark procedures, and performance indicators.
result Identifies mismatches between HWAs and DL models.
Optimal benchmark design varies based on costs in financial manipulation.
problem Manipulation of price benchmarks in finance.
method Analyzes empirical pattern and cost structures to determine optimal benchmark design.
result The optimal benchmark depends on the relative sizes of fixed and variable costs.
Generates synthetic data for benchmarking unsupervised outlier detection.
problem Difficulty in benchmarking unsupervised outlier detection due to rare and varied outliers in real data.
method Proposes a generic process to generate synthetic data with insightful characteristics.
result Demonstrates practicality of the generic process through a benchmark with state-of-the-art detection methods.
Paper introduces benchmark-neutral pricing for long-term contracts.
problem High prices of long-term contracts under risk-neutral pricing.
method Uses growth optimal portfolio as numeraire and new pricing measure.
result Identifies minimal possible prices for contingent claims.
In this report, we present a new reinforcement learning (RL) benchmark based on the Sonic the Hedgehog (TM) video game franchise. This benchmark is intended to measure the performance of transfer learning and few-shot learning algorithms in the RL domain. We also present and evaluate some baseline algorithms on the new…
Fidel-TS creates a new benchmark for time series forecasting models.
problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.
This paper benchmarks neural network robustness to corruptions and perturbations.
problem Establishing benchmarks for image classifier robustness to corruptions and perturbations.
method Developed ImageNet-C and ImageNet-P datasets to evaluate robustness to corruptions and perturbations, not adversarial attacks.
result There are negligible changes in relative corruption robustness from AlexNet to ResNet classifiers.
We solve the multi-criteria benchmarking problem by formalizing it as a social choice problem and identifying conditions for meaningful rankings.
problem Aggregating multiple metrics into a single ranking for models in benchmarking problems.
method Formalizing multi-criteria benchmarking as a social choice problem and identifying sufficient conditions for meaningful rankings.
result We prove that meaningful multi-criteria benchmarking becomes possible under certain preference conditions (single-peaked, group-separable, distance-restricted).
Study proposes new methods to calculate probabilistic benchmarks in noisy data.
problem Identifying opportunities for improvement in comparable units with noisy data.
method 2-step methodology involving undersampling and relevance vector machine.
result Higher discrimination power achieved with macro-economic environment variables.
New framework assesses and benchmarks ML methods for multivariate time series.
problem Benchmarking and explaining performance of machine learning methods.
method Proposes a new framework with systematized performance-explainability characteristics.
result Illustrates application to multivariate time series classifiers.
We study the pricing and hedging of derivatives in incomplete financial markets by considering the local risk-minimization method in the context of the benchmark approach, which will be called benchmarked local risk-minimization. We show that the proposed benchmarked local risk-minimization allows to handle under extre…
The paper shows that benchmark-neutral pricing minimizes option prices.
problem Pricing extreme-maturity European put options on diversified indices.
method Benchmark-neutral pricing applied to a drifted time-transformed squared Bessel process.
result Benchmark-neutral price is the minimal possible price, risk-neutral price is more expensive.
NAS-Bench-Suite simplifies NAS evaluation across diverse tasks.
problem Limited and inconsistent NAS benchmarks hinder research reproducibility.
method Developed a comprehensive, extensible NAS benchmark suite.
result Many NAS conclusions do not generalize across different benchmarks.
Wiki-CS dataset benchmarks Graph Neural Networks using Wikipedia articles.
problem Benchmarking Graph Neural Networks on a new domain with structural differences.
method Derived from Wikipedia, nodes represent Computer Science articles, edges from hyperlinks, 10 classes for different branches, evaluated semi-supervised node classification and link prediction.
result Graph Neural Networks perform well on Wiki-CS, showing structural differences from earlier benchmarks.
Benchmark study evaluates 8 clustering methods on 99 UCR time series datasets.
problem Assessing the performance of clustering methods on time series data.
method Examines 8 clustering methods across 3 categories and 3 distance measures on 99 UCR datasets.
result Provides a comprehensive dataset-level assessment of clustering methods.
BAT benchmark for autobidding tasks in RTB auctions.
problem Lack of comprehensive datasets and benchmarks for autobidding.
method Developed a benchmark for two auction formats, implemented robust baselines.
result Provides a framework for developing and refining autobidding algorithms.
Randomly guessing weights helps analyze RL benchmarks objectively.
problem Understanding the complexity of reinforcement learning benchmarks.
method Generate policy networks by randomly guessing their parameters, evaluate on benchmarks, and analyze results.
result Small untrained networks can provide a robust baseline for various RL tasks.
Paper proposes a new daily benchmark for post-GFC government bond CIP deviations.
problem Lack of a canonical daily benchmark for CIP deviations.
method Used G10 plus KRW currency-tenor panels to analyze three lagged public state variables.
result Three lagged public state variables deliver strong performance in daily regressions.
Benchmarking recursive collapse claims with a new framework under false-positive control.
problem Evaluating recursive systems for failure patterns and warning claims.
method Developed Loopzero framework for testing recursive failures, specified claim boundaries in Lean, evaluated under FP constraint, and compared with standard detectors.
result No standard detectors or Loopzero's pre-registered quantile detector achieved the required operating point under the false-positive contract.
New meta-score EPP interprets model performance differences.
problem Lack of interpretable benchmarks for model performance.
method Elo-based Predictive Power (EPP) meta-score, logistic regression.
result EPP scores have probabilistic interpretation and can be compared between data sets.
Efficiently predict LLM benchmarks using feature selection and regression.
problem Predicting full benchmark scores with minimal question subsets.
method Multiple regression with feature selection, using kernel ridge regression and mRMR.
result Improved prediction accuracy and ranking correlation across various benchmarks.
Paper creates benchmarks for neural hyperparameter search.
problem Difficulty in comparing HPO methods due to high computational costs.
method Developed benchmarks for a feed forward neural network on four regression datasets.
result Exhaustive comparison of HPO methods on the benchmarks.
Public benchmark for machine learning models in critical care.
problem Lack of public benchmarks for machine learning in critical care.
method Defined four tasks (mortality prediction, length of stay, phenotyping, decompensation risk) and compared clinical and deep learning models on eICU dataset.
result First public benchmark on multi-centre critical care dataset, comparing clinical models with predictive models.
Benchmark proposes to assess molecule docking efficiency.
problem Lack of realistic benchmarks for measuring progress in drug design.
method Proposes a docking-based benchmark using SMINA software.
result Graph-based generative models fail to generate high-scoring molecules.
This work introduces a new benchmark to compare neural network training algorithms.
problem Lack of reliable benchmarks to compare training algorithms effectively.
method Developed a new benchmark called AlgoPerf: Training Algorithms benchmark.
result Demonstrated the feasibility of the benchmark and set a provisional state-of-the-art.
The study initiates a theoretical analysis of dynamic benchmarking models.
problem Lack of theoretical foundation and empirical studies in dynamic benchmarks.
method Examined two realizations of dynamic benchmarking: sequential and hierarchical dependency models.
result Sequential dynamic benchmarks show initial performance improvement but can stall after three rounds due to label noise.
Study benchmarks TSC algorithms in distinguishing diffusions using the likelihood ratio test.
problem Benchmarking optimality of TSC algorithms in distinguishing diffusion processes.
method Proposes to benchmark TSC algorithms using the likelihood ratio test (LRT).
result LRT benchmarks are computationally efficient and can be applied to various time series types.
RealCause provides a realistic benchmark for causal inference.
problem Lack of a reliable benchmark for comparing causal effect estimators.
method Flexible generative models to create a benchmark that is both ground-truth and realistic.
result Evaluation of over 1500 causal estimators provides evidence for choosing hyperparameters using predictive metrics.
Proposes a new framework for optimizing utility with state-dependent benchmarks.
problem Various interpretations of benchmarks in utility functions.
method General framework of state-dependent utility optimization with stochastic benchmarks.
result Provides optimal solutions and addresses issues of well-definedness and feasibility.
Language model benchmarks often misrepresent true understanding, revealing vulnerabilities in evaluation methods.
problem Language model benchmarks fail to accurately reflect true language understanding and adaptability.
method Systematic analysis of NLP evaluation frameworks, identifying vulnerabilities in static benchmarks, human evaluation protocols, and LLM-as-judge frameworks.
result Current evaluation methods are unreliable and need improvement to accurately assess LLM performance.
Paper proposes a new speech representation benchmark and model.
problem Lack of benchmarks for comparing speech representations.
method Unsupervised triplet-loss objective for training a universal non-semantic speech representation.
result Proposed representation outperforms other models on benchmark and transfer learning tasks.