This paper studies GAIL's global convergence for general MDP and nonlinear rewards.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New IRL algorithm identifies optimal reward and policy from expert demonstrations.
A network of spiking agents learns complex tasks using global reward signals.
RWR converges to global optimum in certain settings.
Reinforcement learning for embodied agents is a challenging problem. The accumulated reward to be optimized is often a very rugged function, and gradient methods are impaired by many local optimizers. We demonstrate, in an experimental setting, that incorporating an intrinsic reward can smoothen the optimization landsc…
A scalable MARL algorithm using local rewards for cooperative multi-agent learning.
New method uses LP to achieve optimal sample complexity in multi-agent reinforcement learning.
A new algorithm balances global reward and group constraints in federated multi-armed bandits.
Recently, GAIL framework and various variants have shown remarkable possibilities for solving practical MDP problems. However, detailed researches of low-level, and high-dimensional state input in this framework, such as image sequences, has not been conducted. Furthermore, the cost function learned in the traditional …
New algorithm tackles multi-agent bandits with heavy-tailed data.
Many cooperative multiagent reinforcement learning environments provide agents with a sparse team-based reward, as well as a dense agent-specific reward that incentivizes learning basic skills. Training policies solely on the team-based reward is often difficult due to its sparsity. Furthermore, relying solely on the a…
New framework improves restless bandit policies for large numbers of arms.
New algorithms improve privacy in bandit problems with partial information.
LNUCB-TA improves MAB performance by dynamically adjusting exploration rates and recognizing spatiotemporal patterns.
This work introduces reward teaching for federated multi-armed bandits to guide clients towards global optimality.
Study uses RL to optimize global equity portfolios, finds mixed results.
DSAC improves cooperative MARL with general utilities, converging faster than existing methods.
We consider a multi-armed bandit problem in a setting where each arm produces a noisy reward realization which depends on an observable random covariate. As opposed to the traditional static multi-armed bandit problem, this setting allows for dynamically changing rewards that better describe applications where side inf…
This work tackles risk-sensitive deep RL by optimizing policies with variance constraints.
GAIL with neural networks converges to global optima and has a known rate.
Study shows Bitcoin security tied to mining rewards and prices.
We study the global convergence of generative adversarial imitation learning for linear quadratic regulators, which is posed as minimax optimization. To address the challenges arising from non-convex-concave geometry, we analyze the alternating gradient algorithm and establish its Q-linear rate of convergence to a uniq…
Inverse optimal control, also known as inverse reinforcement learning, is the problem of recovering an unknown reward function in a Markov decision process from expert demonstrations of the optimal policy. We introduce a probabilistic inverse optimal control algorithm that scales gracefully with task dimensionality, an…
Novel evolutionary strategy solves stochastic constrained optimization problems.
This study optimizes offline reinforcement learning methods for various tasks without rewards.
New algorithms improve exploration in unbounded reward settings.
The goal of the inverse reinforcement learning (IRL) problem is to recover the reward functions from expert demonstrations. However, the IRL problem like any ill-posed inverse problem suffers the congenital defect that the policy may be optimal for many reward functions, and expert demonstrations may be optimal for man…
This paper proposes a definition of system health in the context of multiple agents optimizing a joint reward function. We use this definition as a credit assignment term in a policy gradient algorithm to distinguish the contributions of individual agents to the global reward. The health-informed credit assignment is t…
We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on the average vectorial outcome. The problem models applications such as multi-objective optimization, maximum entropy exploration, and constrai…
Improves Bayesian optimization using Gaussian process Thompson sampling.
Reinforcement learning agents need exploratory behaviors to escape from local optima. These behaviors may include both immediate dithering perturbation and temporally consistent exploration. To achieve these, a stochastic policy model that is inherently consistent through a period of time is in desire, especially for t…
We introduce reinforcement learning for heterogeneous teams in which rewards for an agent are additively factored into local costs, stimuli unique to each agent, and global rewards, those shared by all agents in the domain. Motivating domains include coordination of varied robotic platforms, which incur different costs…
New learning methods for open systems with variable agents.
This paper suggests a learning-theoretic perspective on how synaptic plasticity benefits global brain functioning. We introduce a model, the selectron, that (i) arises as the fast time constant limit of leaky integrate-and-fire neurons equipped with spiking timing dependent plasticity (STDP) and (ii) is amenable to the…
We consider the combinatorial multi-armed bandit (CMAB) problem, where the reward function is nonlinear. In this setting, the agent chooses a batch of arms on each round and receives feedback from each arm of the batch. The reward that the agent aims to maximize is a function of the selected arms and their expectations…
Reinforcement learning in multi-agent scenarios is important for real-world applications but presents challenges beyond those seen in single-agent settings. We present an actor-critic algorithm that trains decentralized policies in multi-agent settings, using centrally computed critics that share an attention mechanism…
In this study, we investigate the use of global information to speed up the learning process and increase the cumulative rewards of reinforcement learning (RL) in competition tasks. Within the actor-critic RL, we introduce multiple cooperative critics from two levels of the hierarchy and propose a reinforcement learnin…
New method learns high-quality Laplacian representations for reinforcement learning.
A new algorithm reduces communication costs for collaborative decision-making across clients.
Paper tackles free rider attacks in federated learning.
After the shocking series of bankruptcies started in 2008, the public does not trust anymore the classical methods of assessing business risks. The global economic severe downturn caused demand for both developed and emerging economies' exports to drop and the crisis became truly global. However, this current crisis of…
A new method for MARL with partial observations reduces communication overhead.
EGFs use ergodicity to simplify generative flows for easier training and imitation learning.
DTS improves inference-time alignment of diffusion models with less compute.
In many professons employees are rewarded according to their relative performance. Corresponding economy can be modeled by taking independent agents who gain from the market with a rate which depends on their current gain. We argue that this simple realistic rate generates a scale free distribution even though intr…
The paper tackles non-cumulative objectives in reinforcement learning and proposes modifications to existing algorithms.
We describe a novel algorithm for noisy global optimisation and continuum-armed bandits, with good convergence properties over any continuous reward function having finitely many polynomial maxima. Over such functions, our algorithm achieves square-root regret in bandits, and inverse-square-root error in optimisation, …
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents. Motivated by decentralized applications such as sensor networks, swarm robotics, and power grids, we study policy evaluation in MARL, where agents with jo…