Scaling up model and data size improves imitation learning in single-agent games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper analyzes a class of infinite-time-horizon stochastic games with singular controls motivated from the partially reversible problem. It provides an explicit solution for the mean-field game (MFG) and presents sensitivity analysis to compare the solution for the MFG with that for the single-agent control proble…
RL in MFGs is as hard as solving many single-agent RL problems.
Study on multi-agent decision making complexity, showing sample efficiency gaps.
Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior is an approximately optimal solution to an unknown decision problem. These tech…
New approach tackles non-stationary multi-agent games with black-box methods.
New algorithm finds Nash equilibrium in multi-agent games.
Paper studies constrained control games with a novel approximation method.
Despite the notable successes in video games such as Atari 2600, current AI is yet to defeat human champions in the domain of real-time strategy (RTS) games. One of the reasons is that an RTS game is a multi-agent game, in which single-agent reinforcement learning methods cannot simply be applied because the environmen…
Transformers learn to play games in-context, proving Nash equilibrium.
New assumptions and algorithm solve offline two-player zero-sum Markov games.
ARL makes market makers resilient to adversarial conditions.
Recent years have witnessed significant advances in reinforcement learning (RL), which has registered great success in solving various sequential decision-making problems in machine learning. Most of the successful RL applications, e.g., the games of Go and Poker, robotics, and autonomous driving, involve the participa…
PettingZoo library accelerates multi-agent reinforcement learning research.
New RL method explores environments without rewards, achieving efficient policy generation.
Policy gradient method proves convergence in imperfect-information games.
New algorithm tackles multi-agent reinforcement learning with optimal convergence rate.
Develops variational framework for LQG risk-sensitive MFGs with major-minor interactions.
New algorithm for offline RL with linear approx in MDPs and MGs, nearly optimal.
We study a variation of the minority game. There are N agents. Each has to choose between one of two alternatives everyday, and there is reward to each member of the smaller group. The agents cannot communicate with each other, but try to guess the choice others will make, based only the past history of number of peopl…
Adversarial self-play in two-player games has delivered impressive results when used with reinforcement learning algorithms that combine deep neural networks and tree search. Algorithms like AlphaZero and Expert Iteration learn tabula-rasa, producing highly informative training data on the fly. However, the self-play t…
We study discrete-time mean-field Markov games with infinite numbers of agents where each agent aims to minimize its ergodic cost. We consider the setting where the agents have identical linear state transitions and quadratic cost functions, while the aggregated effect of the agents is captured by the population mean o…
This work tackles robust RL in multi-agent settings, improving sample efficiency.
Recent progress in artificial intelligence through reinforcement learning (RL) has shown great success on increasingly complex single-agent environments and two-player turn-based games. However, the real-world contains multiple agents, each learning and acting independently to cooperate and compete with other agents, a…
Improved model-based reinforcement learning for multi-agent Markov games.
A variation of the Minority Game has been applied to study the timing of promotional actions at retailers in the fast moving consumer goods market. The underlying hypotheses for this work are that price promotions are more effective when fewer than average competitors do a promotion, and that a promotion strategy can b…
DORIS algorithm achieves no-regret learning in Markov games with adversarial opponents.
New algorithm improves sample efficiency for zero-sum Markov games.
A promising approach for teaching artificial agents to use natural language involves using human-in-the-loop training. However, recent work suggests that current machine learning methods are too data inefficient to be trained in this way from scratch. In this paper, we investigate the relationship between two categorie…
Paper introduces metrics for evaluating multi-agent policies using best response dynamics.
Reinforcement Learning (RL) is a learning paradigm concerned with learning to control a system so as to maximize an objective over the long term. This approach to learning has received immense interest in recent times and success manifests itself in the form of human-level performance on games like \textit{Go}. While R…
New RL algorithms reduce costs for single-agent and federated learning.
Algorithm learns robust equilibrium in online Markov games with interactive data.
Simplified feature selection using a single agent with restructured choice strategy.
Agents learn to give rewards to others in a shared learning environment.
Optimal penalties for RECs balance environmental and revenue impacts.
Learning by experience in Multi-Agent Systems (MAS) is a difficult and exciting task, due to the lack of stationarity of the environment, whose dynamics evolves as the population learns. In order to design scalable algorithms for systems with a large population of interacting agents (e.g. swarms), this paper focuses on…
Human players in professional team sports achieve high level coordination by dynamically choosing complementary skills and executing primitive actions to perform these skills. As a step toward creating intelligent agents with this capability for fully cooperative multi-agent settings, we propose a two-level hierarchica…
ABPS improves RL training efficiency by sharing policies and evolving hyper-params.
This paper tackles task offloading in edge computing systems with dynamic interactions.
Regret analysis is challenging in Multi-Agent Reinforcement Learning (MARL) primarily due to the dynamical environments and the decentralized information among agents. We attempt to solve this challenge in the context of decentralized learning in multi-agent linear-quadratic (LQ) dynamical systems. We begin with a simp…
Federated learning for combinatorial multi-agent bandits reduces regret and speeds up with fewer communications.
A Systemic Optimal Risk Transfer Equilibrium (SORTE) was introduced in: "Systemic optimal risk transfer equilibrium", Mathematics and Financial Economics (2021), for the analysis of the equilibrium among financial institutions or in insurance-reinsurance markets. A SORTE conjugates the classical Bühlmann's notion of a …
Study differential privacy in multi-agent RL, achieving efficient and private learning.
We look at how asset exchange models can be mapped to random iterated function systems (IFS) giving new insights into the dynamics of wealth accumulation in such models. In particular, we focus on the "yard-sale" (winner gets a random fraction of the poorer players wealth) and the "theft-and-fraud" (winner gets a rando…
We aim to jointly optimize antenna tilt angle, and vertical and horizontal half-power beamwidths of the macrocells in a heterogeneous cellular network (HetNet). The interactions between the cells, most notably due to their coupled interference render this optimization prohibitively complex. Utilizing a single agent rei…
The reinforcement learning community has made great strides in designing algorithms capable of exceeding human performance on specific tasks. These algorithms are mostly trained one task at the time, each new task requiring to train a brand new agent instance. This means the learning algorithm is general, but each solu…
A decentralized approach for agents to learn and optimize collectively.