New method improves consistency of reinforcement learning performance evaluations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Cramming method evaluates learned policies from contextual bandits efficiently.
IRT improves algorithm evaluation across datasets.
Evaluates change point detection algorithms on real-world data.
Recommendation systems have been integrated into the majority of large online systems to filter and rank information according to user profiles. It thus influences the way users interact with the system and, as a consequence, bias the evaluation of the performance of a recommendation algorithm computed using historical…
New evaluation methods for causal models improve algorithm performance assessment.
Recommendation systems have been integrated into the majority of large online systems to filter and rank information according to user profiles. It thus influences the way users interact with the system and, as a consequence, bias the evaluation of the performance of a recommendation algorithm computed using historical…
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these applications is critical for protecting online user experiences but very challenging due to their "p…
Greedy algorithms which use only function evaluations are applied to convex optimization in a general Banach space . Along with algorithms that use exact evaluations, algorithms with approximate evaluations are treated. A priori upper bounds for the convergence rate of the proposed algorithms are given. These bounds…
Generates multimodal safety-critical scenarios for robustness evaluation of decision-making algorithms.
Do-AIQ framework evaluates AI algorithms' quality using DOE.
Backward exploration reduces sample complexity in policy evaluation.
The paper introduces metrics to evaluate NILM algorithms' performance on unseen buildings.
We present and implement two algorithms for analytic asymptotic evaluation of the marginal likelihood of data given a Bayesian network with hidden nodes. As shown by previous work, this evaluation is particularly hard for latent Bayesian network models, namely networks that include hidden variables, where asymptotic ap…
Study improves off-policy evaluation from non-i.i.d. bandit samples.
This paper evaluates algorithms for classification and outlier detection accuracies in temporal data. We focus on algorithms that train and classify rapidly and can be used for systems that need to incorporate new data regularly. Hence, we compare the accuracy of six fast algorithms using a range of well-known time-ser…
New algorithms improve policy evaluation in reinforcement learning.
Optimizes crowdsourced preference-based subjective evaluation with online learning.
Paper applies NEAT for dynamic credit evaluation using streaming data.
PS-BAX uses posterior sampling to select evaluation points for efficient Bayesian algorithm execution.
A new framework evaluates HTE estimators using relative error.
We model and correct bias in sequential evaluation, improving ranking accuracy.
This study evaluates clustering algorithms on high-dimensional data.
In machine learning, the choice of a learning algorithm that is suitable for the application domain is critical. The performance metric used to compare different algorithms must also reflect the concerns of users in the application domain under consideration. In this work, we propose a novel probability-based performan…
Proposes CLRS benchmark to evaluate algorithmic reasoning.
Study limits of testing algorithms without assumptions, finding key performance bounds.
The paper proposes a technique to speed up evolutionary algorithms by using lower-cost approximations of the objective function.
Paper proposes a unified sparsity-based framework for evaluating algorithmic fairness.
In real-world applications of reinforcement learning (RL), noise from inherent stochasticity of environments is inevitable. However, current policy evaluation algorithms, which plays a key role in many RL algorithms, are either prone to noise or inefficient. To solve this issue, we introduce a novel policy evaluation a…
ExperienceThinking optimizes hyperparameters quickly with smart pruning and knowledge use.
Certified algorithms optimize functions with varying costs, providing error bounds.
This paper provides lower bounds on the convergence rate of Derivative Free Optimization (DFO) with noisy function evaluations, exposing a fundamental and unavoidable gap between the performance of algorithms with access to gradients and those with access to only function evaluations. However, there are situations in w…
Discovery of causal relations from observational data is essential for many disciplines of science and real-world applications. However, unlike other machine learning algorithms, whose development has been greatly fostered by a large amount of available benchmark datasets, causal discovery algorithms are notoriously di…
Much attention has been devoted recently to the development of machine learning algorithms with the goal of improving treatment policies in healthcare. Reinforcement learning (RL) is a sub-field within machine learning that is concerned with learning how to make sequences of decisions so as to optimize long-term effect…
A new algorithm optimizes time-varying functions with non-constant evaluation times.
New ranking algorithms are continually being developed and refined, necessitating the development of efficient methods for evaluating these rankers. Online ranker evaluation focuses on the challenge of efficiently determining, from implicit user feedback, which ranker out of a finite set of rankers is the best. Online …
In reinforcement learning (RL) , one of the key components is policy evaluation, which aims to estimate the value function (i.e., expected long-term accumulated reward) of a policy. With a good policy evaluation method, the RL algorithms will estimate the value function more accurately and find a better policy. When th…
NAS helps find best neural network designs.
Paper uses LightGBM for mobile user credit assessment.
New methods for evaluating and optimizing policies in offline RL with unobserved confounders.
We present the first differentially private algorithms for reinforcement learning, which apply to the task of evaluating a fixed policy. We establish two approaches for achieving differential privacy, provide a theoretical analysis of the privacy and utility of the two algorithms, and show promising results on simple e…
This work benchmarks MARL algorithms in cooperative tasks.
We present a method to stop the evaluation of a decision making process when the result of the full evaluation is obvious. This trait is highly desirable for online margin-based machine learning algorithms where a classifier traditionally evaluates all the features for every example. We observe that some examples are e…
Study compares DSPy teleprompter algorithms for aligning LLM evaluations with human annotations.
Cer-Eval saves LLM evaluation costs while maintaining accuracy.
This paper reviews and proposes a new approach for evaluating internal cluster validation indices.
We present FIESTA, a model selection approach that significantly reduces the computational resources required to reliably identify state-of-the-art performance from large collections of candidate models. Despite being known to produce unreliable comparisons, it is still common practice to compare model evaluations base…
In this paper we demonstrate how genetic algorithms can be used to reverse engineer an evaluation function's parameters for computer chess. Our results show that using an appropriate mentor, we can evolve a program that is on par with top tournament-playing chess programs, outperforming a two-time World Computer Chess …