This review explores ML in predicting team sport outcomes, identifying successful strategies and themes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Tennis is a popular sport worldwide, boasting millions of fans and numerous national and international tournaments. Like many sports, tennis has benefitted from the popularity of rigorous record-keeping of game and player information, as well as the growth of machine learning methods for use in sports analytics. Of par…
A simplified Bayesian approach for online sports rating.
This paper optimizes sports betting strategies using neural networks and portfolio theory.
Study proposes new methods to convert betting odds into accurate probabilities for sports forecasting.
Framework for real-time win probability and player ability in sports.
Inspired by applications in sports where the skill of players or teams competing against each other varies over time, we propose a probabilistic model of pairwise-comparison outcomes that can capture a wide range of time dynamics. We achieve this by replacing the static parameters of a class of popular pairwise-compari…
The study analyzes games and social hierarchies, incorporating luck and depth of competition.
Prediction and modelling of competitive sports outcomes has received much recent attention, especially from the Bayesian statistics and machine learning communities. In the real world setting of outcome prediction, the seminal Élő update still remains, after more than 50 years, a valuable baseline which is difficult to…
Study uses complex networks and machine learning to predict soccer match outcomes.
Bayesian model infers strengths from noisy tennis match outcomes.
Decentralized prediction markets use AMMs to pool and withdraw liquidity, improving financial properties.
We study the relationship between social media output and National Football League (NFL) games, using a dataset containing messages from Twitter and NFL game statistics. Specifically, we consider tweets pertaining to specific teams and games in the NFL season and use them alongside statistical game data to build predic…
Comparison data arises in many important contexts, e.g. shopping, web clicks, or sports competitions. Typically we are given a dataset of comparisons and wish to train a model to make predictions about the outcome of unseen comparisons. In many cases available datasets have relatively few comparisons (e.g. there are on…
Accurately predicting the outcome of sporting events has been a goal for many groups who seek to maximize profit. What makes this challenging is that the outcome of an event can be influenced by many factors that dynamically change across time. Oddsmakers attempt to estimate these factors by using both algorithmic and …
Technology has had an unquestionable impact on the way people watch sports. Along with this technological evolution has come a higher standard to ensure a good viewing experience for the casual sports fan. It can be argued that the pervasion of statistical analysis in sports serves to satiate the fan's desire for detai…
The paper analyzes sports commentary to automatically recognize events and extract insights.
Inefficient markets allow investors to consistently outperform the market. To demonstrate that inefficiencies exist in sports betting markets, we created a betting algorithm that generates above market returns for the NFL, NBA, NCAAF, NCAAB, and WNBA betting markets. To formulate our betting strategy, we collected and …
BBE simulates sports betting exchanges for data generation.
Investigates sports betting strategies using modern portfolio theory and Kelly criterion.
Twitter has been proven to be a notable source for predictive modelling on various domains such as the stock market, the dissemination of diseases or sports outcomes. However, such a study has not been conducted in football (soccer) so far. The purpose of this research was to study whether data mined from Twitter can b…
We propose an original model for inferring team strengths using a Markov Random Field, which can be used to generate historical estimates of the offensive and defensive strengths of a team over time. This model was designed to be applied to sports such as soccer or hockey, in which contest outcomes take value in a limi…
Wearables like smartwatches which are embedded with sensors and powerful processors, provide a strong platform for development of analytics solutions in sports domain. To analyze players' games, while motion sensor based shot detection has been extensively studied in sports like Tennis, Golf, Baseball; Table Tennis and…
Paper simplifies complex sports analytics models for better understanding.
Blockchain fan tokens boost sports fan engagement by 50%.
Paper learns skill distributions from game outcomes, proving minimax optimality.
Improved trajectory prediction for team sports using sparse outputs.
The results of data mining endeavors are majorly driven by data quality. Throughout these deployments, serious show-stopper problems are still unresolved, such as: data collection ambiguities, data imbalance, hidden biases in data, the lack of domain information, and data incompleteness. This paper is based on the prem…
New methods for skill rating in sports using state-space models.
This short note is intended as a "Letter to the Editor" Perspective in order that it serves as a contribution, in view of reaching the physics community caring about rare events and scaling laws and unexpected findings, on a domain of wide interest: sport and money. It is apparent from the data reported and discussed b…
BBE simulates betting exchanges to generate synthetic data for AI research.
Machine learning predicts US will win most Olympic medals in 2020.
In-game win probability models, which provide a sports team's likelihood of winning at each point in a game based on historical observations, are becoming increasingly popular. In baseball, basketball and American football, they have become important tools to enhance fan experience, to evaluate in-game decision-making,…
Sporting events are extremely complex and require a multitude of metrics to accurate describe the event. When making multiple predictions, one should make them from a single source to keep consistency across the predictions. We present a multi-task learning method of generating multiple predictions for analysis via a s…
A method for dynamic ranking using BTL model and nearest neighbor rank centrality.
Transportation systems can be conceptualized as an instrument of spreading people and resources over the territory, playing an important role in developing sustainable cities. The current rationale of transport provision is based on population demand, disregarding land use and socioeconomic information. To meet the cha…
The availability of massive data about sports activities offers nowadays the opportunity to quantify the relation between performance and success. In this study, we analyze more than 6,000 games and 10 million events in six European leagues and investigate this relation in soccer competitions. We discover that a team's…
Hierarchical MARL learns complementary skills for team coordination.
Bayesian rating system for large competitions improves prediction and efficiency.
We discuss a possible solution to an unintended consequence of having grades, certificates, rankings and other diversions in the act of transferring knowledge; and zoom in specifically to the topic of having grades, on a curve. We conduct a thought experiment, taking a chapter (and some more?) from the financial market…
The paper sorts big data by revealed preferences, improving consumer and policy decisions.
The paper proposes using experts' insights in machine learning tasks.
Develops a method to infer partial rankings from sparse comparisons.
We generalize Mallows model to learn distance metrics from data.
Study shows news from various topics impacts Nifty 50 index.
Bayesian taut splines estimate modes in probability densities.
Work in Counterfactual Explanations tends to focus on the principle of "the closest possible world" that identifies small changes leading to the desired outcome. In this paper we argue that while this approach might initially seem intuitively appealing it exhibits shortcomings not addressed in the current literature. F…
The paper introduces a new method to find meaningful data subsets in multivariate probability density functions.