We propose a method to efficiently learn diverse strategies in reinforcement learning for query reformulation in the tasks of document retrieval and question answering. In the proposed framework an agent consists of multiple specialized sub-agents and a meta-agent that learns to aggregate the answers from sub-agents to…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DiCE uses diverse agents to explore and learn, avoiding local minima.
Advances in Deep Reinforcement Learning have led to agents that perform well across a variety of sensory-motor domains. In this work, we study the setting in which an agent must learn to generate programs for diverse scenes conditioned on a given symbolic instruction. Final goals are specified to our agent via images o…
InvestorBench benchmarks LLM agents in financial tasks.
Heterogeneous SVO leads to diverse policies in sequential social dilemmas.
This note will extend the research presented in Brown & Rogers (2009) to the case of CRRA agents. We consider the model outlined in that paper in which agents had diverse beliefs about the dividends produced by a risky asset. We now assume that the agents all have CRRA utility, with some integer coefficient of relative…
Agent learns diverse hierarchical structures in unknown environments.
A novel framework uses goal-conditioned reinforcement learning to generate diverse samples.
Unified policy controls diverse agents through modular neural networks.
DIVA generates diverse tasks for complex simulators, enabling adaptive agent training.
In an adaptive population which models financial markets and distributed control, we consider how the dynamics depends on the diversity of the agents' initial preferences of strategies. When the diversity decreases, more agents tend to adapt their strategies together. This change in the environment results in dynamical…
The paper improves QD policy ensembles using distribution ratio estimators.
Adaptive populations such as those in financial markets and distributed control can be modeled by the Minority Game. We consider how their dynamics depends on the agents' initial preferences of strategies, when the agents use linear or quadratic payoff functions to evaluate their strategies. We find that the fluctuatio…
This paper proposes a new method to connect language and physical actions in reinforcement learning.
In this paper we aim to find a measure for the diversity of cash flows between agents in an economy. We argue that cash flows can be linked to probabilities of finding a currency unit in a given cash flow. We then use the information entropy as a natural measure of diversity. This leads to a hirarchical inequality meas…
PCL tackles collaborative learning for diverse agents, reducing sample complexity.
Improves exploration in reinforcement learning with diverse population.
We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally achieves the best action and memory loss that leads to randomized behavior. We show tha…
Self-play is an unsupervised training procedure which enables the reinforcement learning agents to explore the environment without requiring any external rewards. We augment the self-play setting by providing an external memory where the agent can store experience from the previous tasks. This enables the agent to come…
Framework improves agent's ability to learn from noisy images.
Algorithm improves learning by integrating diverse agents' behaviors.
Social learning can make financial markets inefficient, but individual learning can fix this.
KVCOMM optimizes multi-agent LLM systems by reusing KV-caches, reducing redundant processing.
This paper presents a general framework for studying diverse beliefs in dynamic economies. Within this general framework, the characterization of a central-planner general equilbrium turns out to be very easy to derive, and leads to a range of interesting applications. We show how for an economy with log investors hold…
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow. Recent work demonstrated that using a memory buffer of previous successful trajectories can result in more effective…
In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy. During learning, the agent first collects the trajectories into a replay buffer, and later these trajectories are selected randomly for replay. However, the achieved goals in the replay buffer are often biase…
Mixreg improves RL generalization by mixing diverse training environments.
We consider a scenario where an agent has multiple available strategies to explore an unknown environment. For each new interaction with the environment, the agent must select which exploration strategy to use. We provide a new strategy-agnostic method that treat the situation as a Multi-Armed Bandits problem where the…
A method for a single policy to solve various tasks across diverse agent morphologies.
Multi-agent models have been used in many contexts to study generic collective behavior. Similarly, complex networks have become very popular because of the diversity of growth rules giving rise to scale-free behavior. Here we study adaptive networks where the agents trade ``wealth'' when they are linked together while…
New simulation model predicts financial market dynamics with high accuracy.
PEAR dynamically reconfigures agent roles to prevent persistent biases in multi-agent debates.
We propose MAD-GAN, an intuitive generalization to the Generative Adversarial Networks (GANs) and its conditional variants to address the well known problem of mode collapse. First, MAD-GAN is a multi-agent GAN architecture incorporating multiple generators and one discriminator. Second, to enforce that different gener…
Study optimal stopping for group with diverse discount rates using an attitude function.
We have studied numerically the statistical mechanics of the dynamic phenomena, including money circulation and economic mobility, in some transfer models. The models on which our investigations were performed are the basic model proposed by A. Dragulescu and V. Yakovenko [1], the model with uniform saving rate develop…
We use the Minority Game as a testing frame for the problem of the emergence of diversity in socio-economic systems. For the MG with heterogeneous impacts, we show that the direct generalization of the usual agents' profit does not fit some real-world situations. As a typical example we use the traffic formulation of t…
MARL algorithm uses regularization to avoid explicit structures, improving performance.
AutoStan improves Bayesian models via predictive feedback.
This work benchmarks MARL algorithms in cooperative tasks.
We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents' actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions t…
PettingZoo library accelerates multi-agent reinforcement learning research.
Quantitative finance has had a long tradition of a bottom-up approach to complex systems inference via multi-agent systems (MAS). These statistical tools are based on modelling agents trading via a centralised order book, in order to emulate complex and diverse market phenomena. These past financial models have all rel…
We propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing cooperative tasks in partially-observable environments. This targeting behavior is learnt solely from downstream task-specific reward withou…
Solving tasks with sparse rewards is one of the most important challenges in reinforcement learning. In the single-agent setting, this challenge is addressed by introducing intrinsic rewards that motivate agents to explore unseen regions of their state spaces; however, applying these techniques naively to the multi-age…
Fairness is essential for human society, contributing to stability and productivity. Similarly, fairness is also the key for many multi-agent systems. Taking fairness into multi-agent learning could help multi-agent systems become both efficient and stable. However, learning efficiency and fairness simultaneously is a …
Study optimal investment decisions for diverse risk-tolerant agents.
AutoFS combines trainers to improve feature selection efficiency and effectiveness.
Framework analyzes RL agents' behavior to explain their strengths and weaknesses.