Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

295786114 · Jun 202019922001200920182026
48 results for self-interested agents

Study of repeated principal-agent bandit game with self-interested and exploratory learning agents.

problem Interaction between principal and agent in unknown environments with learning and exploration behaviors.
method Developed algorithms for self-interested and exploratory learning agents with bandit feedback, achieving regret bounds.
result Achieved O~(T2/3)\widetilde{O}(T^{2/3}) regret bound for exploratory learning agent in i.i.d. reward setup.

R2-B2 optimizes game interactions with recursive reasoning.

problem Optimizing interactions between boundedly rational agents with unknown payoff functions.
method Recursive Reasoning-Based Bayesian Optimization (R2-B2) for repeated games.
result R2-B2 achieves faster asymptotic convergence to no regret than non-recursive methods.

M3^3RL trains a manager to infer worker minds and assign tasks for optimal collaboration.

problem Optimal coordination among self-interested agents with diverse preferences and skills.
method Mind-aware Multi-agent Management Reinforcement Learning (M^3RL) that infers worker minds and assigns tasks.
result Effective in modeling worker minds and achieving optimal ad-hoc teaming.

A fair reward system boosts participation in federated learning.

problem Fairness in federated learning among competitive agents with siloed data.
method Hierarchically fair federated learning (HFFL) framework with proportional rewards based on contribution levels.
result Efficacy of HFFL in maintaining fairness and facilitating federated learning in competitive settings.

We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally achieves the best action and memory loss that leads to randomized behavior. We show tha…

2004-08-20abs ↗pdf ↗

Supply chain resilience depends on balancing competition and self-interest.

problem Maintaining resilience in decentralized supply chains under uncertainty and competition.
method Modeling competitive suppliers and retailers with yield uncertainty and congestion, analyzing network formation.
result Decentralized supply chains can form resilient networks through competition and self-interest, contrary to intuition.

A decentralized approach for agents to learn and optimize collectively.

problem Challenges in coordinating non-cooperative agents to solve complex sequential decision problems.
method Designing a learning environment where agents learn by trading and optimizing local objectives, leading to a Nash equilibrium.
result Decentralized reinforcement learning algorithms that can handle various decision-making scenarios.

This paper addresses reward estimation and incentive design for agents with hidden rewards.

problem Estimating and incentivizing agents with unknown rewards in a learning setting.
method Repeated adverse selection game with a self-interested learning agent and a learning principal. Introduces an estimator for consistent reward estimation and a data-driven incentive policy.
result Finite-sample consistency of the estimator and a rigorous regret bound for the principal.

This paper tackles sample elicitation for learning systems, introducing a method to incentivize truthful samples.

problem Eliciting credible training samples for complex distributions from humans is challenging.
method Introduces a deep learning aided method to incentivize truthful samples from self-interested and rational agents.
result Achieves approximate incentive compatibility in eliciting truthful samples via accurate estimation of ff-divergence function.

New measure of policy regret shows compatibility with traditional external regret in adversarial games.

problem Incompatibility between traditional and new policy regret measures in adaptive adversaries.
method Revisited policy regret and compared it with external regret; introduced policy equilibrium.
result Policy regret and external regret are compatible in adversarial games.

Paper proposes incentive mechanism to encourage participation in federated learning.

problem Users are reluctant to participate in federated learning due to privacy concerns.
method Formulated as a two-stage Stackelberg game, designed an incentive mechanism to select and compensate users.
result Demonstrated effectiveness of the proposed incentive mechanism through simulations.

The key characteristic of a true free market economy is that exchanges are entirely voluntary. When there is a monopoly in the creation of currency as we have in today's markets, you no longer have a true free market. Features of the current economic system such as central banking and taxation would be nonexistent in a…

2015-06-12abs ↗pdf ↗

Paper proposes incentives for federated learning to ensure truthful contributions.

problem Ensuring truthful contributions from decentralized users in federated learning.
method Introduces a scoring rule based framework to incentivize truthful reporting of local hypotheses at a Bayesian Nash Equilibrium.
result Proposed solution verified using MNIST and CIFAR-10 datasets, showing decreasing scores for low-quality hypotheses.

Researchers disrupt Gaussian model inference to test adversarial attacks.

problem Disrupting conditional inference in multivariate Gaussian models under adversarial conditions.
method Considered white- and grey-box settings with complete and incomplete knowledge of the Gaussian distribution, respectively. Reduced to quadratic and stochastic quadratic programs. Derived structural properties for solution methods.
result Demonstrated the impact and efficacy of attacks in various applications, including real estate evaluation, interest rate estimation, and signals processing.

Two-stage mechanism designs reduce regret in recommender systems with stochastic covariates.

problem Designing effective recommender systems with user covariates sampled online.
method Two-stage algorithm integrating incentivized exploration with offline learning methods.
result Achieves sublinear regret while maintaining incentive compatibility.

Method models other agents' behaviors without requiring direct observation.

problem Understanding and interacting effectively with other agents in reinforcement learning.
method Extracts representations from local observations of the controlled agent using encoder-decoder architectures.
result The method achieves higher returns than baseline methods in multi-agent environments.

Agent-to-agent finance aims to manage payments and trust for AI agents.

problem Managing financial interactions between autonomous AI agents.
method Develops agent-to-agent finance concept and explores blockchain solutions.
result Agent-to-agent finance can address coordination frictions in financial markets.

AI agents manage portfolios, improving on human oversight.

problem Improving strategic asset allocation for institutional investors.
method 50 specialized agents produce capital market assumptions, construct portfolios, critique, and vote on each other's output.
result Meta-agent compares forecasts with realized returns and improves agent performance.

Classic bandit algorithms are robust to strategic manipulation as long as the total budget is small compared to the time horizon.

problem Behavior of stochastic bandit algorithms under strategic manipulation by self-interested arms.
method Analysis of three popular bandit algorithms: UCB, ε-Greedy, and Thompson Sampling.
result Regret upper bound of O(max{B, KlnT}) for all three algorithms under arbitrary adaptive manipulation.

New algorithm reduces learning regret in multi-agent systems with unknown dynamics.

problem Challenges in decentralized learning due to unknown dynamics and lack of communication.
method Proposed MARL algorithm for two-agent LQ systems with unknown dynamics and one-directional communication.
result Achieved O(T)O(\sqrt{T}) regret bound for multi-agent LQ systems with certain communication patterns.

Agents learn to give rewards to others in a shared learning environment.

problem How to encourage cooperation among RL agents in a shared environment.
method Each agent learns a reward function to influence others, optimizing for its own and others' extrinsic objectives.
result Agents significantly outperform standard RL in Markov games, often finding near-optimal division of labor.

Deep learning agents negotiate contracts with prosocial or selfish behaviors.

problem Training agents to negotiate contracts with varying behaviors.
method Multi-Agent Reinforcement Learning, modeling prosocial and selfish behaviors, training a meta agent.
result Trained agents hold their own against human players and emulate human behavior.

New algorithm reduces regret in multi-agent bandits with malicious agents.

problem Collaboration between honest and malicious agents in multi-armed bandits.
method Dynamic reduction of communication with malicious agents, learning who is malicious.
result Algorithm reduces regret even with a single malicious agent, assuming mm is small compared to KK.

We formulate and analyze a multi-agent model for the evolution of individual and systemic risk in which the local agents interact with each other through a central agent who, in turn, is influenced by the mean field of the local agents. The central agent is stabilized by a bistable potential, the only stabilizing force…

2015-07-29abs ↗pdf ↗

Interactive agent modeling by learning to probe improves understanding of other agents' behaviors.

problem Understanding and predicting the behaviors of other agents in interactive scenarios.
method An interactive agent modeling scheme enabled by encouraging the agent to learn to probe, combining imitation learning and curiosity-driven reinforcement learning.
result The agent model learned by the proposed approach generalizes better and enhances performance in multiple applications.

Algorithm maximizes total reward in multi-agent bandits with adversarial corruptions.

problem Maximizing total reward in multi-agent bandits with adversarial corruptions.
method Proposes a cooperative learning algorithm robust to adversarial corruptions.
result Demonstrates an additive O((L/Lmin)C)O((L / L_{\min}) C) regret term for an adversary with unknown corruption budget.

I2C enables agents to learn efficient communication without redundancy.

problem Redundant broadcast communication in multi-agent cooperation.
method I2C learns a prior for agent-agent communication via causal inference and reinforcement learning.
result I2C reduces communication overhead and improves multi-agent cooperative performance.

Modeling and learning turn-taking behaviors in multi-agent systems.

problem Modeling and predicting turn-taking behaviors in dynamic multi-agent systems.
method Individual behavior models (WFSTs) and multi-agent fusion model (logistic regression classifier).
result Accurately models and predicts turn-taking behaviors with high precision.

PEAR dynamically reconfigures agent roles to prevent persistent biases in multi-agent debates.

problem Persistent positional biases and sensitivity to role assignments in fixed topologies.
method Dynamic reconfiguration of agent roles and sparse topologies based on evolving agent states.
result Significantly improves average accuracy over debate baselines across multiple reasoning benchmarks.
Agents Play Mix-gamephysics.soc-ph

In mix-game which is an extension of minority game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. This paper studies the change of the average winnings of agents and volatilities vs. the change of mixture of agents in mix-game model. It finds that the correlatio…

2005-05-17abs ↗pdf ↗

Enactive learning shows agents can learn from their environment, but limited by action choices.

problem Learning and interaction of autonomous agents in complex environments.
method Simulation of artificial agents in maze environments, comparing enactive learning to classical reinforcement learning.
result Enactive agents can learn to avoid unfavorable interactions but performance is limited by action choices.

Many learning agents impact a financial market model, showing complex dynamics.

problem Understanding the dynamics of financial markets with multiple learning agents.
method Agent-based model of financial market with multiple reinforcement learning agents interacting.
result Inclusion of learning agents changes market dynamics to match empirical data.

Agents collaborate to reduce regret in a multi-agent linear bandit problem with side information.

problem Reducing regret in a multi-agent stochastic linear bandit with side information.
method A decentralized algorithm where agents communicate subspace indices and each plays a projected LinUCB on the corresponding low-dimensional subspace.
result Per-agent finite-time regret is much smaller when agents communicate compared to non-communicating case.

A method for efficient reinforcement learning query reformulation.

problem Efficiently learn diverse strategies for query reformulation.
method A framework with specialized sub-agents and a meta-agent trained on full data.
result Improved generalization performance and diversity of reformulation strategies.

Adapts agent strategies on-the-fly for better cross-play in cooperative settings.

problem Cross-play issues between self-play agents and unseen partners.
method Adapts agent strategies using posterior belief updates via Gibbs sampling.
result Achieves strong cross-play in the Hanabi game without prior knowledge of partners' strategies.

Agents are rewarded for influencing others' actions in MARL, improving coordination and communication.

problem Achieving effective coordination and communication in Multi-Agent Reinforcement Learning.
method Rewarding agents for having causal influence over other agents' actions, assessed through counterfactual reasoning.
result Influence rewards lead to enhanced coordination and communication in challenging social dilemma environments.

This paper uses a path integral approach to model complex economic systems with many agents.

problem Modeling economic systems with a large number of interacting agents.
method Develops a path integral formalism to describe the behavior of a large number of agents in an economic system.
result The method provides an analytical treatment of business cycle models with many agents, revealing various phases and interactions.

A novel framework uses goal-conditioned reinforcement learning to generate diverse samples.

problem Generating high-quality, diverse samples from generative models.
method Two agents: GC-agent learns to reconstruct the training set, S-agent learns to imitate GC-agent without knowing the goals.
result Empirically, the method generates diverse and high-quality samples in image synthesis.