Proposes CoPO, a new policy optimization method for competitive games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A drone catches another agile drone using competitive reinforcement learning.
New RL method learns value function for many policies using few key states.
Pakistan examines digital mergers using traditional competition tools.
Dynamic pricing policy converges to Nash equilibrium with low regret.
RL controls small soccer robots in a real league, beating human-designed policies.
We present a broad agenda for meaningful banking regulation reform aiming the creation of evolutive competitive environment to maximize the effectiveness of international financial system through the introduction of fair competition process among the banks in free market capitalism. We assume that the international fin…
Adapts model-based advice to stabilize black-box policies for nonlinear control.
New RL algorithms learn policies competitive with best in class without assuming optimal policy.
This study uses OPE methods to quickly assess auction policies.
New algorithm improves self-play reinforcement learning for competitive games.
Reinforcement learning agents that operate in diverse and complex environments can benefit from the structured decomposition of their behavior. Often, this is addressed in the context of hierarchical reinforcement learning, where the aim is to decompose a policy into lower-level primitives or options, and a higher-leve…
Study model selection in batch policy optimization with three error sources.
Paper tackles online optimization with memory and competitive control.
Study minimax off-policy evaluation in multi-armed bandits with known and unknown behavior policies.
The real estate is a pillar industry of China's national economy. Due to changes in policy and market conditions, the real estate companies are facing greater pressures to survive in a competitive environment. They must improve their financial competitiveness. Based on the conceptual framework of financial competitiven…
MAMBA learns policies competitive with multiple conflicting oracles.
We propose a new algorithm, Mean Actor-Critic (MAC), for discrete-action continuous-state reinforcement learning. MAC is a policy gradient algorithm that uses the agent's explicit representation of all action values to estimate the gradient of the policy, rather than using only the actions that were actually executed. …
Method selects best estimator for off-policy evaluation.
Optimal learning for parametric prophet inequalities with exponential-type distributions
A new imitation learning method uses random search for simple policies, outperforming complex models.
Recent successful deep reinforcement learning algorithms, such as Trust Region Policy Optimization (TRPO) or Proximal Policy Optimization (PPO), are fundamentally variations of conservative policy iteration (CPI). These algorithms iterate policy evaluation followed by a softened policy improvement step. As so, they are…
This paper evaluates financial competitiveness of Indian real estate companies using entropy method.
Greedy policy maximizes information in unknown linear systems.
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are com…
We propose a method for tackling catastrophic forgetting in deep reinforcement learning that is \textit{agnostic} to the timescale of changes in the distribution of experiences, does not require knowledge of task boundaries, and can adapt in \textit{continuously} changing environments. In our \textit{policy consolidati…
Generative neural nets learn deep policies conditioned on goals.
Improved reinforcement learning in Minecraft with human demonstrations.
Reinforcement learning is a promising approach to learning robotics controllers. It has recently been shown that algorithms based on finite-difference estimates of the policy gradient are competitive with algorithms based on the policy gradient theorem. We propose a theoretical framework for understanding this phenomen…
Study optimal treatment assignment policies under strategic agent responses.
A new method solves bilevel optimization problems in competitive Markov games.
We propose a novel approach to train a multi-modal policy from mixed demonstrations without their behavior labels. We develop a method to discover the latent factors of variation in the demonstrations. Specifically, our method is based on the variational autoencoder with a categorical latent variable. The encoder infer…
New findings show optimization is crucial for OPL in large action spaces.
Bayesian design improves by reducing policy training cost.
Learning to cooperate with friends and compete with foes is a key component of multi-agent reinforcement learning. Typically to do so, one requires access to either a model of or interaction with the other agent(s). Here we show how to learn effective strategies for cooperation and competition in an asymmetric informat…
The Pommerman simulation was recently developed to mimic the classic Japanese game Bomberman, and focuses on competitive gameplay in a multi-agent setting. We focus on the 22 team version of Pommerman, developed for a competition at NeurIPS 2018. Our methodology involves training an agent initially through imit…
This paper proposes Self-Imitation Learning (SIL), a simple off-policy actor-critic algorithm that learns to reproduce the agent's past good decisions. This algorithm is designed to verify our hypothesis that exploiting past good experiences can indirectly drive deep exploration. Our empirical results show that SIL sig…
A framework for reinforcement learning tackles CVRP with competitive results.
A new method removes policy optimization in adversarial imitation learning.
In this study, we investigate the use of global information to speed up the learning process and increase the cumulative rewards of reinforcement learning (RL) in competition tasks. Within the actor-critic RL, we introduce multiple cooperative critics from two levels of the hierarchy and propose a reinforcement learnin…
Model-based reinforcement learning has the potential to be more sample efficient than model-free approaches. However, existing model-based methods are vulnerable to model bias, which leads to poor generalization and asymptotic performance compared to model-free counterparts. In addition, they are typically based on the…
A commonly expressed concern about the rise of the peer-to-peer rental market Airbnb is that hosts---those renting out their properties---impose costs on their unwitting neighbors. I consider the question of whether apartment building owners will, in a competitive rental market, set a building-specific Airbnb hosting p…
ZOSPI improves RL policies with global value function exploitation.
RRPI improves offline RL by optimizing policies against worst-case dynamics.
This paper introduces Meta-Q-Learning (MQL), a new off-policy algorithm for meta-Reinforcement Learning (meta-RL). MQL builds upon three simple ideas. First, we show that Q-learning is competitive with state-of-the-art meta-RL algorithms if given access to a context variable that is a representation of the past traject…
PEOC uses policy entropy to detect untrained states in RL.
As the most successful variant and improvement for Trust Region Policy Optimization (TRPO), proximal policy optimization (PPO) has been widely applied across various domains with several advantages: efficient data utilization, easy implementation, and good parallelism. In this paper, a first-order gradient reinforcemen…
Most previous studies on multi-agent reinforcement learning focus on deriving decentralized and cooperative policies to maximize a common reward and rarely consider the transferability of trained policies to new tasks. This prevents such policies from being applied to more complex multi-agent tasks. To resolve these li…