Study introduces a benchmark suite for evaluating neural MI estimators on real-world unstructured datasets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces OARF benchmark suite for federated learning systems.
Progress in machine learning is measured by careful evaluation on problems of outstanding common interest. However, the proliferation of benchmark suites and environments, adversarial attacks, and other complications has diluted the basic evaluation model by overwhelming researchers with choices. Deliberate or accident…
Hyperparameter optimization and neural architecture search can become prohibitively expensive for regular black-box Bayesian optimization because the training and evaluation of a single model can easily take several hours. To overcome this, we introduce a comprehensive tool suite for effective multi-fidelity Bayesian o…
FinTSBridge evaluates financial time series models for asset pricing.
RL Unplugged benchmarks offline RL methods across diverse domains.
Semi-analytic models are best suited to compare galaxy formation and evolution theories with observations. These models rely heavily on halo merger trees, and their realistic features (i.e., no drastic changes on halo mass or jumps on physical locations). Our aim is to provide a new framework for halo merger tree gener…
Study accelerates NAS research with a large dataset of ZC proxies.
LLM Pro Finance Suite enhances financial NLP with instruction-tuned models.
We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE method, leading to …
FAQ efficiently evaluates LLMs with statistical guarantees using adaptive query selection.
FLBench automates federated learning benchmarking.
NAS-Bench-Suite simplifies NAS evaluation across diverse tasks.
We present Simitate --- a hybrid benchmarking suite targeting the evaluation of approaches for imitation learning. A dataset containing 1938 sequences where humans perform daily activities in a realistic environment is presented. The dataset is strongly coupled with an integration into a simulator. RGB and depth stream…
FinSurvival provides a large-scale financial survival modeling benchmark.
Evaluation of deep reinforcement learning (RL) is inherently challenging. In particular, learned policies are largely opaque, and hypotheses about the behavior of deep RL agents are difficult to test in black-box environments. Considerable effort has gone into addressing opacity, but almost no effort has been devoted t…
InvestorBench benchmarks LLM agents in financial tasks.
Equations track profits and losses in trading algorithms.
Deep RL evaluation underestimates uncertainty, leading to misleading conclusions.
YAHPO Gym introduces a new benchmark for evaluating hyperparameter optimization methods.
Open-FinLLMs tackle financial tasks with multimodal capabilities.
Unified evaluation framework for sampling methods.
GLCB uses Gated Linear Networks for online contextual bandits.
Paper establishes baselines for offline RL from visual observations.
The (contextual) multi-armed bandit problem (MAB) provides a formalization of sequential decision-making which has many applications. However, validly evaluating MAB policies is challenging; we either resort to simulations which inherently include debatable assumptions, or we resort to expensive field trials. Recently …
Machine learning research depends on objectively interpretable, comparable, and reproducible algorithm benchmarks. We advocate the use of curated, comprehensive suites of machine learning tasks to standardize the setup, execution, and reporting of benchmarks. We enable this through software tools that help to create an…
Bayesian Optimisation (BO) refers to a suite of techniques for global optimisation of expensive black box functions, which use introspective Bayesian models of the function to efficiently search for the optimum. While BO has been applied successfully in many applications, modern optimisation tasks usher in new challeng…
Knowledge bases contribute to many web search and mining tasks, yet they are often incomplete. To add missing facts to a given knowledge base, various embedding models have been proposed in the recent literature. Perhaps surprisingly, relatively simple models with limited expressiveness often performed remarkably well …
This paper demonstrates the use of genetic algorithms for evolving: 1) a grandmaster-level evaluation function, and 2) a search mechanism for a chess program, the parameter values of which are initialized randomly. The evaluation function of the program is evolved by learning from databases of (human) grandmaster games…
This paper introduces the Behaviour Suite for Reinforcement Learning, or bsuite for short. bsuite is a collection of carefully-designed experiments that investigate core capabilities of reinforcement learning (RL) agents with two objectives. First, to collect clear, informative and scalable problems that capture key is…
One of the primary challenges of visual storytelling is developing techniques that can maintain the context of the story over long event sequences to generate human-like stories. In this paper, we propose a hierarchical deep learning architecture based on encoder-decoder networks to address this problem. To better help…
Study logical generalization in GNNs using a new benchmark.
HEAR benchmark evaluates audio representations for diverse tasks.
This paper describes a method to obtain state model parameters for an infinite series of Links-Gould link invariants LG^{m,n}, based on quantum R matrices associated with the (\dot{0}_m | \dotα_n) representations of the quantum superalgebras U_q[gl(m|n)]. Explicit details of the state models for the cases n=1 and m=1,2…
We design and analyse variations of the classical Thompson sampling (TS) procedure for Bayesian optimisation (BO) in settings where function evaluations are expensive, but can be performed in parallel. Our theoretical analysis shows that a direct application of the sequential Thompson sampling algorithm in either synch…
Reinforcement learning is well suited for optimizing policies of recommender systems. Current solutions mostly focus on model-free approaches, which require frequent interactions with the real environment, and thus are expensive in model learning. Offline evaluation methods, such as importance sampling, can alleviate s…
Scientists have used many different classification methods to solve the problem of music classification. But the efficiency of each classification is different. In this paper, we propose two compared methods on the task of music style classification. More specifically, feature extraction for representing timbral textur…
IRT improves algorithm evaluation across datasets.
New RL environments help AI learn causal relationships from visual data.
GraphBench creates a unified benchmark for graph learning tasks.
Machine learning (ML) has become a vital part in many aspects of our daily life. However, building well performing machine learning applications requires highly specialized data scientists and domain experts. Automated machine learning (AutoML) aims to reduce the demand for data scientists by enabling domain experts to…
Indices of acceptability are well suited to frame the axiomatic features of many performance measures, associated to terminal random cash flows.We extend this notion to classes of càdlàg processes modelling cash flows over a fixed investment horizon.We provide a representation result for bounded paths. We suggest an ac…
SOO uses bandit theory to optimize functions with limited evaluations.
The paper improves off-policy evaluation in contextual bandits using conformal prediction.
Exploiting dependencies between labels is considered to be crucial for multi-label classification. Rules are able to expose label dependencies such as implications, subsumptions or exclusions in a human-comprehensible and interpretable manner. However, the induction of rules with multiple labels in the head is particul…
New scoring rules compare probabilistic top lists in classification.
DIGEN benchmark provides synthetic datasets for ML algorithm evaluation.
We present evidence that the best model for empirical volume-price distributions is not always the same and it strongly depends in (i) the region of the volume-price spectrum that one wants to model and (ii) the period in time that is being modelled. To show these two features we analyze stocks of the New York stock ma…