Framework benchmarks optimizers on multiple criteria.
problem Benchmarking optimizers across diverse test functions.
method Union-free generic depth function for partial orders/rankings.
result Identifies central and outlying rankings of optimizers.
New benchmarks for RNA 3D structure-function modeling.
problem Lack of standardized benchmarks for RNA deep learning.
method Developed seven benchmark datasets, provided tools for data handling, and offered a user-friendly environment for model comparison.
result Demonstrated utility with baseline results using a relational graph neural network.
New benchmark for EEG-eye movement reconstruction from functional data.
problem Reconstructing eye movements from EEG data.
method Functional neural networks and open challenges for evaluation.
result Baseline results for consumer-grade and research-grade hardware.
DIGEN benchmark provides synthetic datasets for ML algorithm evaluation.
problem Understanding and comparing machine learning algorithms' performance.
method Synthetic datasets generated using 40 mathematical functions to evaluate machine learning algorithms.
result DIGEN resource facilitates understanding why algorithms perform poorly and provides ideas for improvement.
Study constructs a Japanese financial LLM benchmark.
problem Need for domain-specific benchmarks for LLMs.
method Constructed a benchmark with multiple Japanese and financial domain tasks.
result GPT-4 outperforms other models in the benchmark.
Combines absolute and relative wealth in portfolio optimization with power utility functions.
problem Optimizing portfolios with both absolute and relative wealth considerations.
method Integrates power utility functions for absolute and relative wealth, considering multiple benchmarks.
result Obtains an explicit solution for portfolio optimization combining absolute and relative wealth.
Proposes a new framework for optimizing utility with state-dependent benchmarks.
problem Various interpretations of benchmarks in utility functions.
method General framework of state-dependent utility optimization with stochastic benchmarks.
result Provides optimal solutions and addresses issues of well-definedness and feasibility.
Natural experiment dataset reveals inconsistent treatment effect estimators.
problem Inconsistent results from over 20 estimators on a new dataset.
method Created a benchmark to evaluate estimator accuracy, derived variance formula, introduced new estimator.
result Doubly robust estimators outperform others by orders of magnitude.
This study provides benchmarks for different implementations of LSTM units between the deep learning frameworks PyTorch, TensorFlow, Lasagne and Keras. The comparison includes cuDNN LSTMs, fused LSTM variants and less optimized, but more flexible LSTM implementations. The benchmarks reflect two typical scenarios for au…
This study benchmarks anomaly detection methods in a functional setup.
problem Efficient anomaly detection in high-frequency sensor data.
method Functional analysis approach for multivariate data.
result Comparison of anomaly detection methods on real datasets.
Time series classification problems have drawn increasing attention in the machine learning and statistical community. Closely related is the field of functional data analysis (FDA): it refers to the range of problems that deal with the analysis of data that is continuously indexed over some domain. While often employi…
Survey and benchmark high-dimensional Bayesian optimization of discrete sequences.
problem Heterogeneous experimental set-ups and technical barriers in high-dimensional Bayesian optimization of discrete sequences.
method Unified framework and software libraries to test and benchmark methods.
result Unified framework and software libraries for testing and benchmarking high-dimensional Bayesian optimization methods.
GPU-accelerated particle methods outperform neural samplers in LFT benchmarks.
problem High-dimensional multimodal sampling problems in lattice field theory.
method GPU-accelerated particle Monte Carlo methods (Sequential Monte Carlo and nested sampling).
result These methods match or outperform neural samplers in sample quality and wall-clock time.
The study examines how permutation-based optimization performance varies across different function representations.
problem Understanding how the order of function evaluations affects optimization performance.
method Iterative search setting with sampling without replacement, algebraic function recombination, correlation analysis, hierarchical clustering, PCA, ANOVA.
result Algebraically modified benchmarks yield stable re-rankings and coherent clusters of functions and sampling policies, indicating non-additive search effort.
The paper introduces a new divergence for portfolio management to outperform a benchmark.
problem Maximizing expected utility of outperformance over a benchmark with constraints.
method Uses α-Bregman-Wasserstein divergence to penalize underperformance more than overperformance. result Proves existence and uniqueness of optimal portfolio strategy and conditions for constraints binding.
Semiparametric method removes bias in functional bilevel gradient estimation.
problem First-order bias in plug-in hypergradient when lower-level problem is nonparametric.
method Semiparametric debiasing theory based on efficient influence function leads to cross-fitted orthogonal hypergradient estimator.
result Asymptotic normality and uniform control over outer parameter established for the estimator.
Study benchmarks label noise detection methods, identifying best practices.
problem Label noise in real-world datasets affects model performance and evaluation reliability.
method Decomposed detection methods into label agreement, aggregation, and information gathering components; introduced a unified benchmark task and novel metric.
result In-sample probability aggregation with logit margin label agreement function achieves best results across scenarios.
Informer model with GMADL loss outperforms benchmarks in high frequency Bitcoin trading.
problem Developing automated trading strategies for high frequency Bitcoin data.
method Informer architecture with RMSE, GMADL, and Quantile loss functions.
result Informer model with GMADL loss function outperforms benchmarks in trading outcomes.
The optimization of algorithm (hyper-)parameters is crucial for achieving peak performance across a wide range of domains, ranging from deep neural networks to solvers for hard combinatorial problems. The resulting algorithm configuration (AC) problem has attracted much attention from the machine learning community. Ho…
Benchmark proposes to assess molecule docking efficiency.
problem Lack of realistic benchmarks for measuring progress in drug design.
method Proposes a docking-based benchmark using SMINA software.
result Graph-based generative models fail to generate high-scoring molecules.
Motivated by the asset-liability management of a nuclear power plant operator, we consider the problem of finding the least expensive portfolio, which outperforms a given set of stochastic benchmarks. For a specified loss function, the expected shortfall with respect to each of the benchmarks weighted by this loss func…
We introduce COCO, an open source platform for Comparing Continuous Optimizers in a black-box setting. COCO aims at automatizing the tedious and repetitive task of benchmarking numerical optimization algorithms to the greatest possible extent. The platform and the underlying methodology allow to benchmark in the same f…
Method for creating synthetic multi-fidelity data sets.
problem Lack of representative synthetic datasets for multifidelity optimisation benchmarks.
method Systematic generation of synthetic fidelities from preexisting datasets.
result Allows systematic investigation of lower fidelity proxies' influence.
In stochastic portfolio theory, a relative arbitrage is an equity portfolio which is guaranteed to outperform a benchmark portfolio over a finite horizon. When the market is diverse and sufficiently volatile, and the benchmark is the market or a buy-and-hold portfolio, functionally generated portfolios introduced by Fe…
The paper clarifies conditions for using benchmark scores in machine learning.
problem Using benchmark scores to draw scientific inferences about learning problems.
method Developing conditions of construct validity inspired by psychological measurement theory.
result Clarifies conditions under which benchmark scores support diverse scientific claims.
PPG separates policy and value function training phases for better reinforcement learning efficiency.
problem Challenges in traditional reinforcement learning methods for policy and value function optimization.
method Integrates Phasic Policy Gradient framework that splits policy and value function training into distinct phases.
result Significantly improves sample efficiency on Procgen Benchmark compared to PPO.
A new machine learning model uses matrix exponentials for universal approximation.
problem Developing a robust and efficient machine learning model.
method Introduces a novel architecture using matrix exponentials as the only nonlinearity.
result The model achieves universal approximation properties and outperforms other models on benchmark tasks.
Develops deep learning methods for solving S-shaped utility maximisation problems.
problem Optimizing portfolios with S-shaped utility and random benchmarks.
method Uses deep learning and duality methods to solve the Hamilton-Jacobi-Bellman equation and adjoint equation.
result Demonstrates the accuracy of deep learning methods for non-concave utility maximisation problems.
The paper extends Merton's problem by adding benchmark tracking, finding optimal strategies.
problem Maximizing consumption utility with a trade-off against benchmark performance.
method Developed a convex duality theorem and derived optimal strategies for specific cases.
result Found optimal portfolio and consumption strategies for CRRA utility and geometric Brownian motion benchmarks.
A natural p-classes generalization of the eXclusive OR problem, the subtraction modulo p, where p is prime, is presented and solved using a single fully connected hidden layer with p-neurons. Although the problem is very simple, the landscape is intricate and challenging and represents an interesting benchmark for grad…
Study benchmarks methods for learning non-Cartesian k-space trajectories and reconstruction.
problem Benchmarking methods for learning non-Cartesian k-space trajectories and reconstruction.
method Comparing PILOT, BJORK, and HybLearn schemes to learn non-Cartesian k-space trajectories and reconstruction.
result HybLearn scheme outperforms other methods in learning and comparing non-Cartesian k-space trajectories and reconstruction.
Develops interpretable low-dimensional kernels with conic discriminant functions.
problem Improving interpretability in kernel-based classification models.
method Gradually constructs simple feature maps leading to interpretable low-dimensional kernels.
result Obtains high accuracy results without extensive hyperparameter tuning.
Bayesian optimization improves molecule design by addressing three pitfalls.
problem Bayesian optimization pitfalls cause poor performance in molecule design.
method Identified and addressed three pitfalls: incorrect prior width, over-smoothing, and inadequate acquisition function maximization.
result Basic BO setup achieves highest performance on PMO benchmark.
This paper explores different graph neural network functions to improve graph isomorphism.
problem Lack of robust implementation for graph neural networks due to limited analysis of underlying functions.
method Examines various alternative functions for different modules in GNNs using benchmark datasets.
result Generally used underlying techniques do not always capture the overall graph structure.
LOB-Bench benchmarks generative AI for financial data, outperforming traditional models.
problem Lack of consensus on evaluating generative AI models for financial data.
method Python-based benchmark with LOB statistics and market impact metrics.
result Generative autoregressive models outperform traditional models in LOB data.
New benchmark PVR tests neural network reasoning about indirection.
problem Understanding neural network generalization limits.
method Introducing Pointer Value Retrieval (PVR) benchmark.
result Large variations in performance across different conditions.
TailedTS dataset benchmarks heavy-tailed time series forecasting and periodicity quantification.
problem Benchmarking robustness of time series models under heavy-tailed distributions.
method Derived from Wikipedia page views, introduces periodicity quantification and robust loss functions.
result Standard Gaussian models degrade on high-volume page categories, while robust alternatives perform consistently.
Study introduces a benchmark suite for evaluating neural MI estimators on real-world unstructured datasets.
problem Lack of comprehensive evaluation methods for neural MI estimators on real-world unstructured datasets.
method Developed a benchmark suite using same-class sampling and a binary symmetric channel trick.
result Showed accurate manipulation of true MI values of real-world datasets.
New acquisition functions improve Bernoulli LSE.
problem Efficiently estimating regions where a Bernoulli function is above or below a threshold.
method Developed new look-ahead acquisition functions for Gaussian process classification models.
result Demonstrated clear benefits of new acquisition functions on benchmark and real-world tasks.
This work explores function-space inference using KL divergence and proposes Bayesian linear regression as a benchmark.
problem Approximating the predictive posterior distribution of Bayesian models without parameter posterior approximation.
method Employing Kullback-Leibler divergence and proposing featurized Bayesian linear regression as a benchmark.
result Minimizing KL divergence leads to an ill-defined objective function, highlighting limitations of this approach.
We consider multi-task regression models where the observations are assumed to be a linear combination of several latent node functions and weight functions, which are both drawn from Gaussian process priors. Driven by the problem of developing scalable methods for forecasting distributed solar and other renewable powe…
We use surrogate losses to obtain several new regret bounds and new algorithms for contextual bandit learning. Using the ramp loss, we derive new margin-based regret bounds in terms of standard sequential complexity measures of a benchmark class of real-valued regression functions. Using the hinge loss, we derive an ef…
Survey and framework for efficient active learning in structural reliability.
problem Efficiently solving complex structural reliability problems.
method Generalized modular framework combining surrogate model, reliability estimation algorithm, learning function, and stopping criterion.
result 39 strategies for solving 20 reliability benchmark problems, highlighting the importance of surrogates and algorithms.
Study tests whether trade-off functions are above or below benchmarks using finite samples.
problem Testing trade-off functions between unknown distributions.
method Identifies a condition for nontrivial testing, constructs a test with error guarantees, and inverts the test for confidence bands.
result Finite-sample testing is possible under specific structural assumptions about rejection regions.
Quantization based techniques are the current state-of-the-art for scaling maximum inner product search to massive databases. Traditional approaches to quantization aim to minimize the reconstruction error of the database points. Based on the observation that for a given query, the database points that have the largest…
Introduces new performance measures using scaled utility functions.
problem Performance measurement in financial contexts.
method Certainty equivalents defined via scaled utility functions, well-posed portfolio optimization problem under generic conditions.
result Link between portfolio dynamics, benchmark process, and utility function choice in the long-run setting.
This work bridges stochastic interpolants to infinite-dimensional Hilbert spaces.
problem Limited flexibility in generating arbitrary distributions for function-valued data.
method Establishes a rigorous framework for stochastic interpolants in infinite-dimensional Hilbert spaces.
result Achieves state-of-the-art results in conditional generation for complex PDE-based benchmarks.
New depth function for partial orders helps compare machine learning algorithms.
problem Comparing machine learning algorithms using non-standard data types.
method Adapted simplicial depth to partial orders, using ufg depth for comparison.
result Demonstrates promising variety of analysis approaches based on ufg methods.