Extracts StarCraft II tournament data for AI and ML studies.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
StarCraft II poses a grand challenge for reinforcement learning. The main difficulties of it include huge state and action space and a long-time horizon. In this paper, we investigate a hierarchical reinforcement learning approach for StarCraft II. The hierarchy involves two levels of abstraction. One is the macro-acti…
The paper formalizes and analyzes multi-agent Q-learning with value factorization.
A new multi-agent learning method improves performance in complex games.
EMIX minimizes surprise in multi-agent reinforcement learning.
New method combines value function decomposition and policy gradients for cooperative multi-agent reinforcement learning.
QPLEX learns efficient multi-agent Q-values by enforcing IGM principle.
QTRAN++ improves MARL performance in complex environments.
In the last few years, deep multi-agent reinforcement learning (RL) has become a highly active area of research. A particularly challenging class of problems in this area is partially observable, cooperative, multi-agent learning, in which teams of agents must learn to coordinate their behaviour while conditioning only…
Injecting human knowledge is an effective way to accelerate reinforcement learning (RL). However, these methods are underexplored. This paper presents our discovery that an abstract forward model (thought-game (TG)) combined with transfer learning (TL) is an effective way. We take StarCraft II as our study environment.…
DefogGAN predicts hidden RTS game information to aid strategic decision-making.
CollaQ improves multi-agent performance in StarCraft by 40% with fewer samples.
We introduce an approach for deep reinforcement learning (RL) that improves upon the efficiency, generalization capacity, and interpretability of conventional approaches through structured perception and relational reasoning. It uses self-attention to iteratively reason about the relations between entities in a scene a…
Multi-agent reinforcement learning (MARL) has recently received considerable attention due to its applicability to a wide range of real-world applications. However, achieving efficient communication among agents has always been an overarching problem in MARL. In this work, we propose Variance Based Control (VBC), a sim…
We consider the problem of high-level strategy selection in the adversarial setting of real-time strategy games from a reinforcement learning perspective, where taking an action corresponds to switching to the respective strategy. Here, a good strategy successfully counters the opponent's current and possible future st…
The development of human-robot systems able to leverage the strengths of both humans and their robotic counterparts has been greatly sought after because of the foreseen, broad-ranging impact across industry and research. We believe the true potential of these systems cannot be reached unless the robot is able to act w…
We propose and address a novel few-shot RL problem, where a task is characterized by a subtask graph which describes a set of subtasks and their dependencies that are unknown to the agent. The agent needs to quickly adapt to the task over few episodes during adaptation phase to maximize the return in the test phase. In…
Study shows Elo models fail to accurately measure transitive strength in competitive games.
Prevalent theories in cognitive science propose that humans understand and represent the knowledge of the world through causal relationships. In making sense of the world, we build causal models in our mind to encode cause-effect relations of events and use these to explain why new events happen. In this paper, we use …
Adversarial training, a special case of multi-objective optimization, is an increasingly prevalent machine learning technique: some of its most notable applications include GAN-based generative modeling and self-play techniques in reinforcement learning which have been applied to complex games such as Go or Poker. In p…
In many real-world settings, a team of agents must coordinate their behaviour while acting in a decentralised way. At the same time, it is often possible to train the agents in a centralised fashion in a simulated or laboratory setting, where global state information is available and communication constraints are lifte…
Recent advances in deep generative models have lead to remarkable progress in synthesizing high quality images. Following their successful application in image processing and representation learning, an important next step is to consider videos. Learning generative models of video is a much harder task, requiring a mod…
The paper reveals a spinning top geometry in real-world games.
RODE learns roles to simplify multi-agent tasks.
LICA learns credit assignment for cooperative agents without explicit formulation.
QR-MIX models joint state-action values as a distribution to handle randomness in MARL.
REFIL learns from imagined sub-group interactions to improve multi-agent reinforcement learning.
Image classification with deep neural networks is typically restricted to images of small dimensionality such as 224 x 244 in Resnet models [24]. This limitation excludes the 4000 x 3000 dimensional images that are taken by modern smartphone cameras and smart devices. In this work, we aim to mitigate the prohibitive in…
Flexible decentralized MARL framework for cooperative multi-agent learning.
FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.
A novel memory mechanism for reinforcement learning agents that stores past events in human-readable language.
In complex tasks, such as those with large combinatorial action spaces, random exploration may be too inefficient to achieve meaningful learning progress. In this work, we use a curriculum of progressively growing action spaces to accelerate learning. We assume the environment is out of our control, but that the agent …
Framework improves self-play for cooperative multi-agent learning.
Learning when to communicate and doing that effectively is essential in multi-agent tasks. Recent works show that continuous communication allows efficient training with back-propagation in multi-agent scenarios, but have been restricted to fully-cooperative tasks. In this paper, we present Individualized Controlled Co…
NeuPL learns diverse policies in strategy games efficiently.
Deep reinforcement learning has been successful in a variety of tasks, such as game playing and robotic manipulation. However, attempting to learn \textit{tabula rasa} disregards the logical structure of many domains as well as the wealth of readily available knowledge from domain experts that could help "warm start" t…
Deep RL improves Diplomacy performance, outperforming previous methods.
Improved exploration in cooperative multi-agent reinforcement learning.
Reinforcement learning encounters major challenges in multi-agent settings, such as scalability and non-stationarity. Recently, value function factorization learning emerges as a promising way to address these challenges in collaborative multi-agent systems. However, existing methods have been focusing on learning full…
Effective coordination is crucial to solve multi-agent collaborative (MAC) problems. While centralized reinforcement learning methods can optimally solve small MAC instances, they do not scale to large problems and they fail to generalize to scenarios different from those seen during training. In this paper, we conside…
In this study, we investigated the (K,H), (K,K_{II}), (H,K_{II})-Weingarten and (K,H),(K,K_{II}),(H,K_{II}) and (K,H,K_{II})-linear Weingarten canal surfaces in IR^3.
QMIX combines per-agent values to create decentralised policies.
In this study, we analyze the general canal surfaces in terms of the features flat, II-flat minimality and II-minimality, namely we study under which conditions the first and second Gauss and mean curvature vanishes, i.e. K=0, H=0, K_{II}=0 and H_{II} =0. We give a non-existence result for general canal surfaces in E^3…
Strict type-II blowup in harmonic map flow is proven to have Hölder continuous body map.
Equivariant MuZero improves generalization in procedurally generated environments.
Paper proves uniqueness of Type II Yamabe metrics on manifolds.
Ozawa solution describes surface deformation from Davey-Stewartson II equation.
In this paper, we study stability and instability problem for type-II partitioning problem. First, we make a complete classification of stable type-II stationary hypersurfaces in a ball in a space form as totally geodesic -balls. Second, for general ambient spaces and convex domains, we give some topological restric…