New algorithm for multi-player bandits with selfish players, achieving logarithmic regret.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm outperforms existing ones in multi-player bandit problems without sensing.
We present an effective technique for training deep learning agents capable of negotiating on a set of clauses in a contract agreement using a simple communication protocol. We use Multi Agent Reinforcement Learning to train both agents simultaneously as they negotiate with each other in the training environment. We al…
Multi-player Multi-Armed Bandits (MAB) have been extensively studied in the literature, motivated by applications to Cognitive Radio systems. Driven by such applications as well, we motivate the introduction of several levels of feedback for multi-player MAB algorithms. Most existing work assume that sensing informatio…
This paper analyzes the profitability of selfish mining on blockchain, considering the risk of ruin.
We calculate the dynamics of tax evasion within a multi-agent econophysics model which is adopted from the theory of magnetism and previously has been shown to capture the main characteristics from agent-based based models which build on the standard Allingham and Sandmo approach. In particular, we implement a feedback…
Human behavioural patterns exhibit selfish or competitive, as well as selfless or altruistic tendencies, both of which have demonstrable effects on human social and economic activity. In behavioural economics, such effects have traditionally been illustrated experimentally via simple games like the dictator and ultimat…
New algorithms for batch decision-making with high-dimensional user data.
Standard economic theory, starting with Adam Smith's invisible hand, holds that those who trade for their own selfish motives of maximizing their private preferences may contribute more to the public wealth than those who claim altruistic motives. Under restrictive conditions, this has been shown to result from a self-…
We study stochastic multi-armed bandits with many players. The players do not know the number of players, cannot communicate with each other and if multiple players select a common arm they collide and none of them receive any reward. We consider the static scenario, where the number of players remains fixed, and the d…
We consider a symmetric multi-players zero-sum game with two strategic variables. There are players, . Each player is denoted by . Two strategic variables are and , . They are related by invertible functions. Using the minimax theorem by \cite{sion} we will show that Nas…
A multi-player bandit system resists adversarial attacks with near-optimal regret.
We consider two-player non-zero-sum stopping games in discrete time. Unlike Dynkin games, in our games the payoff of each player is revealed after both players stop. Moreover, each player can adjust her own stopping strategy according to the other player's action. In the first part of the paper, we consider the game wh…
We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, , the reward distributions for a given arm are heterog…
New algorithms for n-player games using a player-centered approach.
We develop a machine learning approach to represent and analyze the underlying spatial structure that governs shot selection among professional basketball players in the NBA. Typically, NBA players are discussed and compared in an heuristic, imprecise manner that relies on unmeasured intuitions about player behavior. T…
New algorithm tackles multi-player bandit problems with limited access to arms.
Paper presents Transfer Portal model for accurate player performance predictions.
A new algorithm RESYNC for defenders against malicious attackers in multi-player bandits.
Study predicts soccer player market values using machine learning and SHAP for interpretability.
Algorithm optimizes multi-player learning with noisy rewards without direct communication.
New strategy achieves optimal regret without communication or collisions in multi-player bandit.
New algorithm for multi-player bandits in decentralized, asynchronous systems.
We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this results in a maximal loss. We prove the first -type regret guarantee for th…
The possibility of using player engagement predictions to profile high spending video game users is explored. In particular, individual-player survival curves in terms of days after first login, game level reached and accumulated playtime are used to classify players into different groups. Lifetime value predictions fo…
In this paper we present an early Apprenticeship Learning approach to mimic the behaviour of different players in a short adaption of the interactive fiction Anchorhead. Our motivation is the need to understand and simulate player behaviour to create systems to aid the design and personalisation of Interactive Narrativ…
New algorithm reduces regret in strategic prediction problem.
Method predicts NBA players' multi-modal movement trajectories.
New learning dynamics adapt to corrupted games, improving performance in real-world scenarios.
Paper presents content-based models for game recommendation in cold start scenarios.
We study a multiplayer stochastic multi-armed bandit problem in which players cannot communicate, and if two or more players pull the same arm, a collision occurs and the involved players receive zero reward. We consider the challenging heterogeneous setting, in which different arms may have different means for differe…
Algorithm reduces regret in multi-player bandits with unknown collision rewards.
Assessing the impact of the individual actions performed by soccer players during games is a crucial aspect of the player recruitment process. Unfortunately, most traditional metrics fall short in addressing this task as they either focus on rare actions like shots and goals alone or fail to account for the context in …
New framework values football players based on in-game interactions.
We consider a setting where multiple players sequentially choose among a common set of actions (arms). Motivated by a cognitive radio networks application, we assume that players incur a loss upon colliding, and that communication between players is not possible. Existing approaches assume that the system is stationary…
The paper analyzes a game where players must balance short-term and long-term interests, leading to cooperative or competitive outcomes.
Researchers predict NBA player salaries using machine learning, avoiding overfitting.
Over the last few decades, the player recruitment process in professional football has evolved into a multi-billion industry and has thus become of vital importance. To gain insights into the general level of their candidate reinforcements, many professional football clubs have access to extensive video footage and adv…
A variety of machine learning models have been proposed to assess the performance of players in professional sports. However, they have only a limited ability to model how player performance depends on the game context. This paper proposes a new approach to capturing game context: we apply Deep Reinforcement Learning (…
MpFL models clients as strategic players to reach equilibrium with less communication.
Algorithm aggregates rewards from multiple players to learn related tasks in online bandit learning.
The in-game economies of massively multi-player online games (MMOGs) are complex systems that have to be carefully designed and managed. This paper presents the results of an analysis of auction house data from the MMOG Glitch, across a 14 month time period, the entire lifetime of the game. The data comprise almost 3 m…
Study proposes new OPE estimators for two-player zero-sum games.
The paper addresses the Multiplayer Multi-Armed Bandit (MMAB) problem, where decision makers or players collaborate to maximize their cumulative reward. When several players select the same arm, a collision occurs and no reward is collected on this arm. Players involved in a collision are informed about this collis…
New algorithms tackle adversarial multi-player bandits with forced-collision communication.
Recent innovations in Information and Communication Technologies (ICT) provide new opportunities and challenges for integration of distributed energy resources (DERs) into the energy supply system as active market players. By increasing integration of DERs, novel market platform should be designed for these new market …
Develops a framework to estimate NBA player salary ROI.
Study of a game with multiple players and common shocks using probabilistic methods.