Study of repeated principal-agent bandit game with self-interested and exploratory learning agents.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recent advances in Bayesian reinforcement learning (BRL) have shown that Bayes-optimality is theoretically achievable by modeling the environment's latent dynamics using Flat-Dirichlet-Multinomial (FDM) prior. In self-interested multi-agent environments, the transition dynamics are mainly controlled by the other agent'…
R2-B2 optimizes game interactions with recursive reasoning.
A multi-agent system is trialed as a means of crowd-sourcing inexpensive but high quality streams of predictions. Each agent is a microservice embodying statistical models and endowed with economic self-interest. The ability to fork and modify simple agents is granted to a large number of employees in a firm and empiri…
A fair reward system boosts participation in federated learning.
We consider a simple model of a closed economic system where the total money is conserved and the number of economic agents is fixed. In analogy to statistical systems in equilibrium, money and the average money per economic agent are equivalent to energy and temperature, respectively. We investigate the effect of the …
We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally achieves the best action and memory loss that leads to randomized behavior. We show tha…
We calculate the dynamics of tax evasion within a multi-agent econophysics model which is adopted from the theory of magnetism and previously has been shown to capture the main characteristics from agent-based based models which build on the standard Allingham and Sandmo approach. In particular, we implement a feedback…
Real-time advertising allows advertisers to bid for each impression for a visiting user. To optimize specific goals such as maximizing revenue and return on investment (ROI) led by ad placements, advertisers not only need to estimate the relevance between the ads and user's interests, but most importantly require a str…
Most of the prior work on multi-agent reinforcement learning (MARL) achieves optimal collaboration by directly controlling the agents to maximize a common reward. In this paper, we aim to address this from a different angle. In particular, we consider scenarios where there are self-interested agents (i.e., worker agent…
A decentralized approach for agents to learn and optimize collectively.
This paper addresses reward estimation and incentive design for agents with hidden rewards.
Supply chains are the backbone of the global economy. Disruptions to them can be costly. Centrally managed supply chains invest in ensuring their resilience. Decentralized supply chains, however, must rely upon the self-interest of their individual components to maintain the resilience of the entire chain. We examine t…
It is important to collect credible training samples for building data-intensive learning systems (e.g., a deep learning system). Asking people to report complex distribution , though theoretically viable, is challenging in practice. This is primarily due to the cognitive loads required for human agents t…
The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such as external regret do not take into account. We revisit the notion of policy regret and first show that there are online learning settings …
Paper proposes incentive mechanism to encourage participation in federated learning.
The key characteristic of a true free market economy is that exchanges are entirely voluntary. When there is a monopoly in the creation of currency as we have in today's markets, you no longer have a true free market. Features of the current economic system such as central banking and taxation would be nonexistent in a…
Paper proposes incentives for federated learning to ensure truthful contributions.
The seniority of debt, which determines the order in which a bankrupt institution repays its debts, is an important and sometimes contentious feature of financial crises, yet its impact on system-wide stability is not well understood. We capture seniority of debt in a multiplex network, a graph of nodes connected by mu…
Motivated by economic applications such as recommender systems, we study the behavior of stochastic bandits algorithms under \emph{strategic behavior} conducted by rational actors, i.e., the arms. Each arm is a \emph{self-interested} strategic player who can modify its own reward whenever pulled, subject to a cross-per…
Researchers disrupt Gaussian model inference to test adversarial attacks.
Two-stage mechanism designs reduce regret in recommender systems with stochastic covariates.
Method models other agents' behaviors without requiring direct observation.
Agent-to-agent finance aims to manage payments and trust for AI agents.
AI agents manage portfolios, improving on human oversight.
RL agents outperform baselines in asset allocation.
Agents learn to give rewards to others in a shared learning environment.
New algorithm reduces regret in multi-agent bandits with malicious agents.
Regret analysis is challenging in Multi-Agent Reinforcement Learning (MARL) primarily due to the dynamical environments and the decentralized information among agents. We attempt to solve this challenge in the context of decentralized learning in multi-agent linear-quadratic (LQ) dynamical systems. We begin with a simp…
We formulate and analyze a multi-agent model for the evolution of individual and systemic risk in which the local agents interact with each other through a central agent who, in turn, is influenced by the mean field of the local agents. The central agent is stabilized by a bistable potential, the only stabilizing force…
Algorithm maximizes total reward in multi-agent bandits with adversarial corruptions.
I2C enables agents to learn efficient communication without redundancy.
We present an effective technique for training deep learning agents capable of negotiating on a set of clauses in a contract agreement using a simple communication protocol. We use Multi Agent Reinforcement Learning to train both agents simultaneously as they negotiate with each other in the training environment. We al…
PEAR dynamically reconfigures agent roles to prevent persistent biases in multi-agent debates.
In mix-game which is an extension of minority game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. This paper studies the change of the average winnings of agents and volatilities vs. the change of mixture of agents in mix-game model. It finds that the correlatio…
We propose a method for modeling and learning turn-taking behaviors for accessing a shared resource. We model the individual behavior for each agent in an interaction and then use a multi-agent fusion model to generate a summary over the expected actions of the group to render the model independent of the number of age…
Many learning agents impact a financial market model, showing complex dynamics.
Reinforcement learning (RL) algorithms allow agents to learn skills and strategies to perform complex tasks without detailed instructions or expensive labelled training examples. That is, RL agents can learn, as we learn. Given the importance of learning in our intelligence, RL has been thought to be one of key compone…
Agents collaborate to reduce regret in a multi-agent linear bandit problem with side information.
Adapts agent strategies on-the-fly for better cross-play in cooperative settings.
A novel framework uses goal-conditioned reinforcement learning to generate diverse samples.
Agents need world models to generalize multi-step tasks.
ODC protocol improves learning in asynchronous multi-agent bandits.
The ability of modeling the other agents, such as understanding their intentions and skills, is essential to an agent's interactions with other agents. Conventional agent modeling relies on passive observation from demonstrations. In this work, we propose an interactive agent modeling scheme enabled by encouraging an a…
The behavioral dynamics of multi-agent systems have a rich and orderly structure, which can be leveraged to understand these systems, and to improve how artificial agents learn to operate in them. Here we introduce Relational Forward Models (RFM) for multi-agent learning, networks that can learn to make accurate predic…
EPC curriculum improves MARL performance as agent population grows.
Agents learn social skills from each other, improving performance.
This paper presents an analytical treatment of economic systems with an arbitrary number of agents that keeps track of the systems' interactions and agents' complexity. This formalism does not seek to aggregate agents. It rather replaces the standard optimization approach by a probabilistic description of both the enti…