New algorithm for multi-player bandits without needing lower bounds or scaling inversely.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Algorithm reduces regret in multi-player bandits with unknown collision rewards.
This paper tackles global Nash equilibrium in non-convex multi-player games.
New strategy achieves optimal regret without communication or collisions in multi-player bandit.
New algorithm tackles multi-player bandit problems with limited access to arms.
A new algorithm reduces regret in multi-player bandits without collision info.
New algorithm outperforms existing ones in multi-player bandit problems without sensing.
A new algorithm RESYNC for defenders against malicious attackers in multi-player bandits.
New algorithms tackle adversarial multi-player bandits with forced-collision communication.
New algorithm for multi-player bandits with collision-dependent rewards.
Multi-player Multi-Armed Bandits (MAB) have been extensively studied in the literature, motivated by applications to Cognitive Radio systems. Driven by such applications as well, we motivate the introduction of several levels of feedback for multi-player MAB algorithms. Most existing work assume that sensing informatio…
Optimistic Thompson Sampling reduces regret in unknown multi-player games.
No communication allows optimal instance-dependent regret guarantees in multi-player bandits.
A multi-player bandit system resists adversarial attacks with near-optimal regret.
We consider a setting where multiple players sequentially choose among a common set of actions (arms). Motivated by a cognitive radio networks application, we assume that players incur a loss upon colliding, and that communication between players is not possible. Existing approaches assume that the system is stationary…
New algorithm for multi-player bandits in decentralized, asynchronous systems.
We study the multi-player stochastic multiarmed bandit (MAB) problem in an abruptly changing environment. We consider a collision model in which a player receives reward at an arm if it is the only player to select the arm. We design two novel algorithms, namely, Round-Robin Sliding-Window Upper Confidence Bound\# (RR-…
We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this results in a maximal loss. We prove the first -type regret guarantee for th…
We consider a symmetric multi-players zero-sum game with two strategic variables. There are players, . Each player is denoted by . Two strategic variables are and , . They are related by invertible functions. Using the minimax theorem by \cite{sion} we will show that Nas…
Paper closes the gap in MP-MAB problems with novel adaptive communication and exploration.
Computing Nash equilibrium (NE) of multi-player games has witnessed renewed interest due to recent advances in generative adversarial networks. However, computing equilibrium efficiently is challenging. To this end, we introduce the Gradient-based Nikaido-Isoda (GNI) function which serves: (i) as a merit function, vani…
The in-game economies of massively multi-player online games (MMOGs) are complex systems that have to be carefully designed and managed. This paper presents the results of an analysis of auction house data from the MMOG Glitch, across a 14 month time period, the entire lifetime of the game. The data comprise almost 3 m…
We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, , the reward distributions for a given arm are heterog…
This paper shows how to learn variational inequalities fast with strong monotonicity.
Motivated by cognitive radios, stochastic multi-player multi-armed bandits gained a lot of interest recently. In this class of problems, several players simultaneously pull arms and encounter a collision - with 0 reward - if some of them pull the same arm at the same time. While the cooperative case where players maxim…
Novel algorithms for multi-agent reinforcement learning reduce sample complexity.
We prove a general connection between the communication complexity of two-player games and the sample complexity of their multi-player locally private analogues. We use this connection to prove sample complexity lower bounds for locally differentially private protocols as straightforward corollaries of results from com…
Study shows how multiple traders can trade together without excessive price impact.
The paper tackles multi-player information asymmetry bandits in metric spaces.
Survey on multiplayer bandits, highlighting theoretical gaps and future directions.
We introduce a new type of graphical model called a "cumulative distribution network" (CDN), which expresses a joint cumulative distribution as a product of local functions. Each local function can be viewed as providing evidence about possible orderings, or rankings, of variables. Interestingly, we find that the condi…
Data-driven modeling increasingly requires to find a Nash equilibrium in multi-player games, e.g. when training GANs. In this paper, we analyse a new extra-gradient method for Nash equilibrium finding, that performs gradient extrapolations and updates on a random subset of players at each iteration. This approach prova…
Cross-dimensional neural networks improve AI in Catan game.
In illiquid markets, option traders may have an incentive to increase their portfolio value by using their impact on the dynamics of the underlying. We provide a mathematical framework within which to value derivatives under market impact in a multi-player framework by introducing strategic interactions into the Almgre…
New learning dynamics adapt to corrupted games, improving performance in real-world scenarios.
We consider a variant of the stochastic multi-armed bandit problem, where multiple players simultaneously choose from the same set of arms and may collide, receiving no reward. This setting has been motivated by problems arising in cognitive radio networks, and is especially challenging under the realistic assumption t…
New algorithms improve dueling bandit performance in multiplayer settings.
Extends trading framework to incorporate real-world constraints.
Motivated by the scarcity of accurate payoff feedback in practical applications of game theory, we examine a class of learning dynamics where players adjust their choices based on past payoff observations that are subject to noise and random disturbances. First, in the single-player case (corresponding to an agent tryi…
We study stochastic multi-armed bandits with many players. The players do not know the number of players, cannot communicate with each other and if multiple players select a common arm they collide and none of them receive any reward. We consider the static scenario, where the number of players remains fixed, and the d…
Dialogue research tends to distinguish between chit-chat and goal-oriented tasks. While the former is arguably more naturalistic and has a wider use of language, the latter has clearer metrics and a straightforward learning signal. Humans effortlessly combine the two, for example engaging in chit-chat with the goal of …
Algorithm aggregates rewards from multiple players to learn related tasks in online bandit learning.
Paper solves learning imperfect-information games with fewer episodes.
Game theory models how agents trade in a risky asset considering price impact and a common signal.
Study proposes new OPE estimators for two-player zero-sum games.
This paper presents the first two editions of Visual Doom AI Competition, held in 2016 and 2017. The challenge was to create bots that compete in a multi-player deathmatch in a first-person shooter (FPS) game, Doom. The bots had to make their decisions based solely on visual information, i.e., a raw screen buffer. To p…
Unified framework for predicting data changes influenced by predictions.
We introduce CSE for MLSF games and devise online learning algorithms for achieving no-external Stackelberg-regret.