Research tackles alliance formation in many-player zero-sum games, showing reinforcement learning fails but a contract mechanism can help.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study stochastic multi-armed bandits with many players. The players do not know the number of players, cannot communicate with each other and if multiple players select a common arm they collide and none of them receive any reward. We consider the static scenario, where the number of players remains fixed, and the d…
Study optimal portfolios for many players in a market model with random coefficients.
Over the last few decades, the player recruitment process in professional football has evolved into a multi-billion industry and has thus become of vital importance. To gain insights into the general level of their candidate reinforcements, many professional football clubs have access to extensive video footage and adv…
New algorithm tackles multi-player bandit problems with limited access to arms.
We consider the problem of learning in single-player and multiplayer multiarmed bandit models. Bandit problems are classes of online learning problems that capture exploration versus exploitation tradeoffs. In a multiarmed bandit model, players can pick among many arms, and each play of an arm generates an i.i.d. rewar…
Almost universally, wealth is not distributed uniformly within societies or economies. Even though wealth data have been collected in various forms for centuries, the origins for the observed wealth-disparity and social inequality are not yet fully understood. Especially the impact and connections of human behavior on …
MpFL models clients as strategic players to reach equilibrium with less communication.
Algorithm aggregates rewards from multiple players to learn related tasks in online bandit learning.
New algorithm for multi-player bandits in decentralized, asynchronous systems.
We propose and study the known-compensation multi-arm bandit (KCMAB) problem, where a system controller offers a set of arms to many short-term players for steps. In each step, one short-term player arrives to the system. Upon arrival, the player aims to select an arm with the current best average reward and receiv…
Study many-player investment-consumption games with power FPPs, finding market-risk preference affects consumption.
Our work extends Coase's theorem to settings with uncertainty, showing how to maximize social welfare through property rights and learning.
The target of -armed bandit problem is to find the global maximum of an unknown stochastic function , given a finite budget of evaluations. Recently, -armed bandits have been widely used in many situations. Many of these applications need to deal with large-scale data sets. To deal with…
We introduce a new virtual environment for simulating a card game known as "Big 2". This is a four-player game of imperfect information with a relatively complicated action space (being allowed to play 1,2,3,4 or 5 card combinations from an initial starting hand of 13 cards). As such it poses a challenge for many curre…
New methods learn correlated equilibria in large games without structural assumptions.
Novel approach finds implicit regularisation in two-player games using BEA.
This paper improves sample efficiency for learning equilibria in multi-player games.
We consider a symmetric multi-players zero-sum game with two strategic variables. There are players, . Each player is denoted by . Two strategic variables are and , . They are related by invertible functions. Using the minimax theorem by \cite{sion} we will show that Nas…
We introduce CSE for MLSF games and devise online learning algorithms for achieving no-external Stackelberg-regret.
A multi-player bandit system resists adversarial attacks with near-optimal regret.
We consider two-player non-zero-sum stopping games in discrete time. Unlike Dynkin games, in our games the payoff of each player is revealed after both players stop. Moreover, each player can adjust her own stopping strategy according to the other player's action. In the first part of the paper, we consider the game wh…
We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, , the reward distributions for a given arm are heterog…
New algorithms for n-player games using a player-centered approach.
New algorithm for multi-player bandits with selfish players, achieving logarithmic regret.
We develop a machine learning approach to represent and analyze the underlying spatial structure that governs shot selection among professional basketball players in the NBA. Typically, NBA players are discussed and compared in an heuristic, imprecise manner that relies on unmeasured intuitions about player behavior. T…
A game-theoretic approach simplifies MBRL design and improves sample efficiency.
Paper presents Transfer Portal model for accurate player performance predictions.
A new algorithm RESYNC for defenders against malicious attackers in multi-player bandits.
Study predicts soccer player market values using machine learning and SHAP for interpretability.
We study the problem of repeated play in a zero-sum game in which the payoff matrix may change, in a possibly adversarial fashion, on each round; we call these Online Matrix Games. Finding the Nash Equilibrium (NE) of a two player zero-sum game is core to many problems in statistics, optimization, and economics, and fo…
Algorithm optimizes multi-player learning with noisy rewards without direct communication.
New strategy achieves optimal regret without communication or collisions in multi-player bandit.
Proposes a flexible tournament design combining knockout and round-robin.
We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this results in a maximal loss. We prove the first -type regret guarantee for th…
The possibility of using player engagement predictions to profile high spending video game users is explored. In particular, individual-player survival curves in terms of days after first login, game level reached and accumulated playtime are used to classify players into different groups. Lifetime value predictions fo…
In this paper we present an early Apprenticeship Learning approach to mimic the behaviour of different players in a short adaption of the interactive fiction Anchorhead. Our motivation is the need to understand and simulate player behaviour to create systems to aid the design and personalisation of Interactive Narrativ…
Deep RL improves Diplomacy performance, outperforming previous methods.
Method predicts NBA players' multi-modal movement trajectories.
New learning dynamics adapt to corrupted games, improving performance in real-world scenarios.
Paper presents content-based models for game recommendation in cold start scenarios.
We study a multiplayer stochastic multi-armed bandit problem in which players cannot communicate, and if two or more players pull the same arm, a collision occurs and the involved players receive zero reward. We consider the challenging heterogeneous setting, in which different arms may have different means for differe…
Algorithm reduces regret in multi-player bandits with unknown collision rewards.
Assessing the impact of the individual actions performed by soccer players during games is a crucial aspect of the player recruitment process. Unfortunately, most traditional metrics fall short in addressing this task as they either focus on rare actions like shots and goals alone or fail to account for the context in …
New framework values football players based on in-game interactions.
We consider a setting where multiple players sequentially choose among a common set of actions (arms). Motivated by a cognitive radio networks application, we assume that players incur a loss upon colliding, and that communication between players is not possible. Existing approaches assume that the system is stationary…
It is usually assumed that stock prices reflect a balance between large numbers of small individual sellers and buyers. However, over the past fifty years mutual funds and other institutional shareholders have assumed an ever increasing part of stock transactions: their assets, as a percentage of GDP, have been multipl…
Game theory models how agents trade in a risky asset considering price impact and a common signal.