Proposes a flexible tournament design combining knockout and round-robin.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Elo ratings learn model parameters quickly using Markov chains.
Classifies intrinsically linked tournaments by their score sequences.
Co-designing efficient machine learning based systems across the whole hardware/software stack to trade off speed, accuracy, energy and costs is becoming extremely complex and time consuming. Researchers often struggle to evaluate and compare different published works across rapidly evolving software frameworks, hetero…
We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between generators and discriminators provide an effective way to evaluate generative models. We introduce two methods for summarizing tournament outcomes…
A directed graph is if every embedding of that graph contains a non-split link , where each component of is a consistently oriented cycle in . A is a directed graph where each pair of vertices is connected by exactly one directed edge. We consider intr…
A regularized risk minimization procedure for regression function estimation is introduced that achieves near optimal accuracy and confidence under general conditions, including heavy-tailed predictor and response variables. The procedure is based on median-of-means tournaments, introduced by the authors in [8]. It is …
Extracts StarCraft II tournament data for AI and ML studies.
Reply to Tetlock et al. on tail risk and probability gap.
Breaks the hardness conjecture for batch RL with a novel tournament-based approach.
We propose a novel ranking model that combines the Bradley-Terry-Luce probability model with a nonnegative matrix factorization framework to model and uncover the presence of latent variables that influence the performance of top tennis players. We derive an efficient, provably convergent, and numerically stable majori…
The paper tackles optimal level set estimation in crowdsourcing and tournaments.
Receiver operating characteristic (ROC) analysis is widely used for evaluating diagnostic systems. Recent studies have shown that estimating an area under ROC curve (AUC) with standard cross-validation methods suffers from a large bias. The leave-pair-out (LPO) cross-validation has been shown to correct this bias. Howe…
The paper derives upper bounds on the MLE error for BTL model under general graphs.
Pairwise comparison data arises in many domains, including tournament rankings, web search, and preference elicitation. Given noisy comparisons of a fixed subset of pairs of items, we study the problem of estimating the underlying comparison probabilities under the assumption of strong stochastic transitivity (SST). We…
We are interested in parallelizing the Least Angle Regression (LARS) algorithm for fitting linear regression models to high-dimensional data. We consider two parallel and communication avoiding versions of the basic LARS algorithm. The two algorithms have different asymptotic costs and practical performance. One offers…
A number of applications (e.g., AI bot tournaments, sports, peer grading, crowdsourcing) use pairwise comparison data and the Bradley-Terry-Luce (BTL) model to evaluate a given collection of items (e.g., bots, teams, students, search results). Past work has shown that under the BTL model, the widely-used maximum-likeli…
New algorithms for model selection in off-policy evaluation of reinforcement learning.
New method improves stock return prediction in non-stationary markets.
Directed graphs occur throughout statistical modeling of networks, and exchangeability is a natural assumption when the ordering of vertices does not matter. There is a deep structural theory for exchangeable undirected graphs, which extends to the directed case via measurable objects known as digraphons. Using digraph…
This paper extends Median-of-Means to new learning problems involving pairwise comparisons.
In this paper we demonstrate how genetic algorithms can be used to reverse engineer an evaluation function's parameters for computer chess. Our results show that using an appropriate mentor, we can evolve a program that is on par with top tournament-playing chess programs, outperforming a two-time World Computer Chess …
The small-ball method was introduced as a way of obtaining a high probability, isomorphic lower bound on the quadratic empirical process, under weak assumptions on the indexing class. The key assumption was that class members satisfy a uniform small-ball estimate: that for given const…
The paper studies dynamic ranking and translation synchronization from evolving pairwise comparison graphs.
In this paper we demonstrate how genetic algorithms can be used to reverse engineer an evaluation function's parameters for computer chess. Our results show that using an appropriate expert (or mentor), we can evolve a program that is on par with top tournament-playing chess programs, outperforming a two-time World Com…
A method for dynamic ranking using BTL model and nearest neighbor rank centrality.
Bayesian model infers strengths from noisy tennis match outcomes.
Recently we proposed a general, ensemble-based feature engineering wrapper (FEW) that was paired with a number of machine learning methods to solve regression problems. Here, we adapt FEW for supervised classification and perform a thorough analysis of fitness and survival methods within this framework. Our tests demon…
Fantasy Premier League (FPL) performance predictors tend to base their algorithms purely on historical statistical data. The main problems with this approach is that external factors such as injuries, managerial decisions and other tournament match statistics can never be factored into the final predictions. In this pa…
New method estimates optimizer for convex stochastic problems.
Tennis is a popular sport worldwide, boasting millions of fans and numerous national and international tournaments. Like many sports, tennis has benefitted from the popularity of rigorous record-keeping of game and player information, as well as the growth of machine learning methods for use in sports analytics. Of par…
We consider the problem of stochastic -armed dueling bandit in the contextual setting, where at each round the learner is presented with a context set of items, each represented by a -dimensional feature vector, and the goal of the learner is to identify the best arm of each context sets. However, unlike the …
Recent progress in artificial intelligence through reinforcement learning (RL) has shown great success on increasingly complex single-agent environments and two-player turn-based games. However, the real-world contains multiple agents, each learning and acting independently to cooperate and compete with other agents, a…
This paper tackles bandit optimization with a new pairwise comparison oracle for unknown strongly concave functions.
Item response theory (IRT) models are widely used in psychometrics and educational measurement, being deployed in many high stakes tests such as the GRE aptitude test. IRT has largely focused on estimation of a single latent trait (e.g. ability) that remains static through the collection of item responses. However, in …
Paper learns skill distributions from game outcomes, proving minimax optimality.
A framework for faster, better infographic design by non-experts and experts alike.
The paper proposes a method to reliably select design algorithms for machine learning-guided design tasks.
Unified approach to experimental design using interlacing polynomials.
We propose a new low-cost machine-learning-based methodology which assists designers in reducing the gap between the problem and the solution in the design process. Our work applies reinforcement learning (RL) to find the optimal task-oriented design solution through the construction of the design action for each task.…
Design-by-Morphing creates radical airfoil designs without geometric constraints.
MO-PaDGAN generates diverse, high-performance designs with multiple metrics.
MCD automates counterfactual design searches for multi-modal tasks.
ANN with GA optimizes flexible disc design for lower mass and stress.
New method for mixed-variable GSA improves material design efficiency.
Deep learning models enhance engineering design automation.
Efficient deep learning computing requires algorithm and hardware co-design to enable specialization: we usually need to change the algorithm to reduce memory footprint and improve energy efficiency. However, the extra degree of freedom from the algorithm makes the design space much larger: it's not only about designin…
DAD learns to design experiments quickly, outperforming traditional methods.