Examines multiagent systems for complex learning tasks.
problem Achieving cohesive learning behavior in multiagent networks.
method General formulation for multiagent dynamics and conditions for learning.
result Conditions for achieving cohesive learning behavior in multiagent networks.
Neural MMO simulates MMOs to study multiagent intelligence.
problem Limited research environments for multiagent intelligence.
method Developed a new game environment inspired by MMOs.
result Standard methods can learn interesting behaviors in MMOs.
Proposes a model to estimate treatment effects in complex multiagent systems over time.
problem Challenges in evaluating interventions in multiagent systems, especially with time-varying relationships and covariates.
method Interpretable counterfactual recurrent network leveraging graph variational recurrent neural networks and domain knowledge.
result Achieved lower estimation errors and more effective treatment timing than baselines in simulated and real-world scenarios.
New MARL method combines agent knowledge to reduce complexity.
problem Curse of dimensionality in multiagent reinforcement learning.
method Decomposes multiagent problem into multi-task problem, uses distillation and value-matching.
result Outperforms policy distillation alone and improves learning in both discrete and continuous action spaces.
New algorithm breaks multiagency gap in robust MARL.
problem Vulnerability of MARL to sim-to-real gaps.
method Distributionally robust Markov games (RMGs) with a new uncertainty set formulation.
result First algorithm to break the curse of multiagency for RMGs.
Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function representing discounted return. In this paper, we examine the role of these policy gradi…
MERL uses evolutionary and gradient-based methods to optimize sparse team-based and dense agent-specific rewards in multiagent coordination.
problem Training multiagent reinforcement learning policies on sparse team-based rewards is difficult and relying solely on agent-specific rewards is sub-optimal.
method MERL employs a split-level training platform with an evolutionary algorithm and a gradient-based optimizer, transferring skills between the two processes.
result MERL significantly outperforms state-of-the-art methods on coordination benchmarks.
Modeling agent behavior is central to understanding the emergence of complex phenomena in multiagent systems. Prior work in agent modeling has largely been task-specific and driven by hand-engineering domain-specific prior knowledge. We propose a general learning framework for modeling agent behavior in any multiagent …
V-learning tackles multiagent reinforcement learning by reducing sample complexity.
problem Curse of multiagents in multiagent reinforcement learning.
method V-learning is a fully decentralized algorithm that learns Nash, correlated, and coarse correlated equilibria.
result V-learning achieves sample complexity that scales with the maximum number of actions per agent, not the joint action space.
Datasets with hundreds to tens of thousands features is the new norm. Feature selection constitutes a central problem in machine learning, where the aim is to derive a representative set of features from which to construct a classification (or prediction) model for a specific task. Our experimental study involves micro…
End-to-end model predicts multiagent trajectories using game theory and neural nets.
problem Predicting trajectories of interacting agents in complex scenarios.
method Hybrid neural net with game-theoretic reasoning, using implicit layers to map preferences to Nash equilibria.
result Trains an interpretable model that predicts future trajectories and transfers to decision making.
In this paper, we predict the likelihood of a player making a shot in basketball from multiagent trajectories. Previous approaches to similar problems center on hand-crafting features to capture domain specific knowledge. Although intuitive, recent work in deep learning has shown this approach is prone to missing impor…
IC3Net improves communication efficiency in multiagent tasks.
problem Efficient communication in multiagent tasks, especially in semi-cooperative and competitive settings.
method IC3Net uses a gating mechanism and individualized rewards to control continuous communication, improving scalability and performance.
result IC3Net yields improved performance and convergence rates compared to baselines as task scale increases.
A game environment simulates competition among many agents for resources.
problem Understanding large-scale multiagent interactions and resource competition.
method Developed a persistent, massively multiplayer AI environment.
result Population size affects the development of skillful behaviors and niche differentiation.
New MARL algorithms resolve the curse of multiagency with function approximation.
problem Challenges in Multi-Agent Reinforcement Learning (MARL) due to the curse of multiagency.
method V-Learning with Policy Replay and Decentralized Optimistic Policy Mirror Descent.
result First polynomial sample complexity results for learning approximate Coarse Correlated Equilibria (CCEs) of Markov Games under decentralized linear function approximation.
MGpi model predicts social actions in group interactions.
problem Social interaction among multiple agents and groups.
method Deep neural network with Kinesic-Proxemic-Message gate for social signal gating.
result Achieves state-of-the-art performance in group identification.
New algorithms for RL in Markov games with independent linear function approximation, breaking the curse of multiagents.
problem Tackles the challenge of learning Markov equilibria in large state space Markov games with multiple agents.
method Proposes independent linear Markov games and designs new algorithms for learning Markov coarse correlated equilibria and Markov correlated equilibria with polynomial sample complexity.
result Breaks the curse of multiagents by achieving sample complexity bounds that scale polynomially with each agent's function class complexity.
A new multiagent model of the stock market is formulated that contains four states in which the agents may be located. Next, the model is reformulated in the language of the functional integral containing fluctuations of prices and quantities of cash flows. It is shown that in the functional integral of that type descr…
Paper shows RLHF can be solved similarly to standard RL.
problem Difficulty of RLHF compared to standard RL.
method Reduction to reward-based RL techniques.
result RLHF can be solved using existing algorithms for reward-based RL.
Neural Replicator Dynamics improves deep RL performance in nonstationary environments.
problem Nonstationarity and instability in multiagent reinforcement learning.
method Derive a new algorithm using replicator dynamics to bypass softmax in policy gradient methods.
result Neural Replicator Dynamics (NeuRD) outperforms policy gradient methods in nonstationary environments.
In this paper it was developed a modification of the known multiagent model Minority Game, designed to simulate the behavior of traders in financial markets and the resulting price dynamics on the abstract resource. The model was implemented in the form of software. The modified version of Minority Game was investigate…
Explains financial market simulation mechanisms and agent behaviors.
problem Necessity of including fundamental value in multiagent financial market simulations.
method Discusses three methods for generating fundamental value and illustrates one Bayesian agent's estimation process.
result Presentation of two widely examined agents: Zero Intelligence and Heuristic Belief Learning.
We discuss the behavior of two magnitudes, physical complexity and mutual information function of the outcome of a model of heterogeneous, inductive rational agents inspired in the El Farol Bar problem and the Minority Game. The first is a measure rooted in Kolmogorov-Chaitin theory and the second one a measure related…
Paper develops a private algorithm for multi-agent learning in bandits.
problem Private cooperative learning in decentralized systems.
method Developed extsc{FedUCB} algorithm for multi-agent learning.
result Improves pseudoregret bounds and empirical performance.
Describes state variables in sequential decision problems, linking them to Markovian and non-Markovian models.
problem Sequential decision problems, especially in active learning and POMDPs, where decisions affect what is observed and learned.
method Canonical framework and novel two-agent perspective of POMDPs, defining state variables to claim Markovian or non-Markovian models.
result Properly modeled sequential decision problems are Markovian, while real decision problems are often non-Markovian.
Deep reinforcement learning has achieved many recent successes, but our understanding of its strengths and limitations is hampered by the lack of rich environments in which we can fully characterize optimal behavior, and correspondingly diagnose individual actions against such a characterization. Here we consider a fam…
Deep RL enhances cyber security through adaptive, responsive, and scalable defenses.
problem Complex, dynamic cyber attacks require adaptive and scalable security solutions.
method Incorporates deep learning into traditional RL for cyber defense.
result Demonstrates the potential of DRL in solving high-dimensional cyber defense problems.
We consider the problem of belief aggregation: given a group of individual agents with probabilistic beliefs over a set of uncertain events, formulate a sensible consensus or aggregate probability distribution over these events. Researchers have proposed many aggregation methods, although on the question of which is be…
Paper develops efficient algorithms for learning rationalizable equilibria in multiplayer games.
problem Learning rationalizable behavior in multiplayer games under bandit feedback.
method New algorithms for finding rationalizable Coarse Correlated Equilibria and Correlated Equilibria with polynomial sample complexity.
result Achieved polynomial sample complexity for learning rationalizable equilibria, improving over existing exponential complexity.
EMIX minimizes surprise in multi-agent reinforcement learning.
problem Surprise and approximation bias in multi-agent reinforcement learning.
method Energy-based MIXer (EMIX) for minimizing surprise across multiple agents.
result EMIX demonstrates consistent stable performance in challenging StarCraft II scenarios.
New method combines value function decomposition and policy gradients for cooperative multi-agent reinforcement learning.
problem Challenges in cooperative multi-agent reinforcement learning, especially credit assignment and large action spaces.
method Decomposed Soft Actor-Critic (mSAC) method with Q network architecture, discrete probabilistic policy, and counterfactual advantage function.
result Significantly outperforms policy-based approach COMA and achieves competitive results with SOTA value-based approach Qmix.
We solve continuous-time reinforcement learning using distributional Hamilton-Jacobi-Bellman equations.
problem Predicting the distribution of returns in continuous-time, stochastic environments.
method We derive a distributional Hamilton-Jacobi-Bellman equation for Itô diffusions and Feller-Dynkin processes, and propose an algorithm based on a JKO scheme.
result We propose an online control algorithm that can be used to approximately solve the distributional HJB equation.
SONATA algorithm converges to solutions of nonconvex smooth functions with KL property.
problem Decentralized optimization over networks with nonconvex smooth functions and convex constraints.
method Decentralized gradient-tracking algorithm SONATA under the KL property.
result SONATA converges to stationary solutions at R-linear rate for θ∈(0,1/2], sublinear rate for θ∈(1/2,1), and R-linear rate for θ=0. MARL improves LBM stability and accuracy across scales.
problem Stability and accuracy issues in under-resolved LBM simulations.
method Multi-Agent Reinforcement Learning (MARL) to dynamically control local relaxation parameters.
result MARL closures stabilize simulations and recover spectra of fully resolved models.
In this paper, we consider a class of finite-sum convex optimization problems defined over a distributed multiagent network with m agents connected to a central server. In particular, the objective function consists of the average of m (≥1) smooth components associated with each network agent together with a s…
BRIDGE improves decentralized learning resilience against Byzantine failures.
problem Decentralized learning in distributed systems is vulnerable to Byzantine failures.
method Introduces BRIDGE, a scalable Byzantine-resilient decentralized gradient descent algorithm.
result Proves algorithmic and statistical convergence guarantees for BRIDGE.