Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Jan 199419922001200920172026
48 results for optimal agent behavior

Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior is an approximately optimal solution to an unknown decision problem. These tech…

2013-08-15abs ↗pdf ↗

FOCOPS optimizes agent's behavior while adhering to constraints.

problem Optimizing agent's behavior while respecting safety constraints.
method FOCOPS solves a constrained optimization problem in policy space, then projects the solution back into the parametric space.
result FOCOPS achieves better performance on constrained robotics tasks.

New algorithm learns satisficing behaviors more efficiently in complex environments.

problem Intractability of optimal exploration in complex environments.
method Extends a deep reinforcement learning agent to learn satisficing policies without model-based planning.
result Demonstrates efficient learning of satisficing behaviors and optimal behaviors when feasible.

We consider model-based reinforcement learning (MBRL) in 2-agent, high-fidelity continuous control problems -- an important domain for robots interacting with other agents in the same workspace. For non-trivial dynamical systems, MBRL typically suffers from accumulating errors. Several recent studies have addressed thi…

2019-01-29abs ↗pdf ↗

We propose a framework for ensuring safe behavior of a reinforcement learning agent when the reward function may be difficult to specify. In order to do this, we rely on the existence of demonstrations from expert policies, and we provide a theoretical framework for the agent to optimize in the space of rewards consist…

2018-05-21abs ↗pdf ↗

Model shows how price impact and transaction costs affect trading behavior and profits.

problem Analyzing trading behavior and profits in markets with transaction costs and price impact.
method Proves the existence of an equilibrium in a model with transaction costs and price impact.
result Existence of a strictly positive optimal transaction cost from the exchange's perspective.

Intrinsically motivated reinforcement learning aims to address the exploration challenge for sparse-reward tasks. However, the study of exploration methods in transition-dependent multi-agent settings is largely absent from the literature. We aim to take a step towards solving this problem. We present two exploration m…

2019-10-12abs ↗pdf ↗

Agents learn to outperform in trading by using past and current prices.

problem Optimal trading performance beyond theoretical limits.
method Two-agent Almgren-Chriss liquidation game, schedule-learning, DDQN architectures.
result Agents with access to past and current prices achieve supra-competitive outcomes.

Generative AI reduces herd behavior in trading, but can also lead to optimal herding.

problem Impact of generative AI on financial stability and herd behavior.
method Laboratory experiments with large language models replicating human trading behavior.
result AI agents make more rational decisions than humans, reducing herd behavior but also potentially leading to optimal herding.

This study benchmarks AI agents for personalized retail promotions using simulations.

problem Optimizing coupon targeting for sparse customer purchase events.
method Comprehensive simulations of customer shopping behaviors; training RL agents on batch data.
result Contextual bandit and deep RL methods outperform static policies in sparse reward environments.

Study optimal investment with herd behavior using rational decision decomposition.

problem Optimal investment problem considering herd behavior between two agents.
method Introduce average deviation term, use variational method, rational decision decomposition, investment opinion.
result Quantitative analysis of herd behavior impact on investment decisions.

Modeling agent behavior is central to understanding the emergence of complex phenomena in multiagent systems. Prior work in agent modeling has largely been task-specific and driven by hand-engineering domain-specific prior knowledge. We propose a general learning framework for modeling agent behavior in any multiagent …

2018-06-17abs ↗pdf ↗

In many sequential decision making tasks, it is challenging to design reward functions that help an RL agent efficiently learn behavior that is considered good by the agent designer. A number of different formulations of the reward-design problem, or close variants thereof, have been proposed in the literature. In this…

2018-04-17abs ↗pdf ↗

The ability of modeling the other agents, such as understanding their intentions and skills, is essential to an agent's interactions with other agents. Conventional agent modeling relies on passive observation from demonstrations. In this work, we propose an interactive agent modeling scheme enabled by encouraging an a…

2018-10-01abs ↗pdf ↗

We propose a method for modeling and learning turn-taking behaviors for accessing a shared resource. We model the individual behavior for each agent in an interaction and then use a multi-agent fusion model to generate a summary over the expected actions of the group to render the model independent of the number of age…

2018-12-10abs ↗pdf ↗

We introduce a strategic behavior in reinsurance bilateral transactions, where agents choose the risk preferences they will appear to have in the transaction. Within a wide class of risk measures, we identify agents' strategic choices to a range of risk aversion coefficients. It is shown that at the strictly beneficial…

2019-09-04abs ↗pdf ↗

Agent-based simulation assesses tradable credit schemes for congestion reduction.

problem Simplistic modeling of TCS impacts in transportation research.
method Agent- and activity-based simulation framework within SimMobility.
result TCS stabilizes network and market performance over time, reducing congestion.

ABIDES-MARL uses MARL to study market behavior in a realistic financial simulation.

problem Understanding equilibrium behavior in complex financial market games.
method Combines MARL with a realistic LOB simulation to study market behavior.
result Validated approach by solving an extended Kyle model and showing how execution strategies shape market dynamics.

Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.

problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.

QuantAgent learns trading signals through self-improvement.

problem Building domain-specific knowledge for LLMs in quantitative investment.
method Two-layer loop approach: inner loop refines responses, outer loop tests and learns.
result QuantAgent approximates optimal trading behavior with provable efficiency.

The paper addresses human-like decision-making in multi-agent systems using bounded risk-sensitive Markov Games.

problem Modeling human-like decision-making in multi-agent systems with risk-seeking and loss-aversion behaviors.
method Forward policy design and inverse reward learning with iterative reasoning and cumulative prospect theory.
result The proposed algorithms demonstrate both risk-averse and risk-seeking behaviors in multi-agent systems.

Study models human investors' sub-rational behavior in financial markets.

problem Lack of a comprehensive model for human sub-rationality in financial markets.
method Flexible reinforcement learning model incorporating five human sub-rational aspects.
result Model accurately reproduces human behavior and reveals insights into market dynamics.

This paper introduces a new approach to active inference using constrained Bethe Free Energy.

problem Tackling the limitations of existing epistemic behavior models in active inference.
method Introducing a constrained Bethe Free Energy (CBFE) perspective to optimize epistemic behavior in generative models.
result CBFE optimization leads to more robust and flexible epistemic behavior compared to existing methods.

A standard belief on emerging collective behavior is that it emerges from simple individual rules. Most of the mathematical research on such collective behavior starts from imperative individual rules, like always go to the center. But how could an (optimal) individual rule emerge during a short period within the group…

2018-02-21abs ↗pdf ↗

The study shows how probability weighting can lead to betting in a risk-averse economy.

problem Understanding how probability weighting affects economic behavior and risk aversion.
method Examining a von Neumann-Morgenstern economy with an RDU agent to model probability weighting effects.
result Probability weighting can lead to endogenous betting in an economy with common beliefs.

Study examines if LLMs' trading styles match real market behavior.

problem Lack of behavioral consistency in LLMs' trading strategies.
method Year-long simulations with LLMs, operationalizing behavioral finance drivers, and comparing with financial theory.
result LLMs' strategy switching is only partially consistent with behavioral finance theories.

A major bottleneck for developing general reinforcement learning agents is determining rewards that will yield desirable behaviors under various circumstances. We introduce a general mechanism for automatically specifying meaningful behaviors from raw pixels. In particular, we train a generative adversarial network to …

2017-11-21abs ↗pdf ↗

Agent-based modeling is a paradigm of modeling dynamic systems of interacting agents that are individually governed by specified behavioral rules. Training a model of such agents to produce an emergent behavior by specification of the emergent (as opposed to agent) behavior is easier from a demonstration perspective. W…

2019-10-10abs ↗pdf ↗

Extends driving model to control agent behavior in simulations.

problem Simulate realistic driving behavior for autonomous systems.
method Introduces Control-ITRA method to influence agent behavior through waypoint assignment and target speed modulation.
result Demonstrates controllable, infraction-free trajectories while preserving realism.

Algorithm learns optimal coordination for strategic agents in uncertain settings.

problem Optimizing rewards for strategic agents with private types and actions.
method Combines delaying mechanism, reward angle estimation, and LinUCB algorithm.
result Near optimal regret bound of O~(T)\tilde{O}(\sqrt{T}) for learning optimal policy.

Optimal risk sharing found for heterogeneous risk attitudes using distortion risk measures.

problem Risk sharing in economies with diverse risk attitudes.
method Modeling preferences with distortion risk measures, using comonotonic and counter-monotonic principles.
result Optimal risk sharing strategies identified based on risk attitudes, reducing the nn-agent problem to a two-agent formulation.

Collective behavior of the complex socio-economic systems is heavily influenced by the herding, group, behavior of individuals. The importance of the herding behavior may enable the control of the collective behavior of the individuals. In this contribution we consider a simple agent-based herding model modified to inc…

2013-09-24abs ↗pdf ↗

Many machine learning problems can be formulated as consensus optimization problems which can be solved efficiently via a cooperative multi-agent system. However, the agents in the system can be unreliable due to a variety of reasons: noise, faults and attacks. Providing erroneous updates leads the optimization process…

2017-10-14abs ↗pdf ↗

Study collaborative learning among multi-agents in multi-armed bandits.

problem Minimizing group cumulative regret in a heterogeneous multi-agent setting.
method Developed decentralized algorithms for collaboration between NN agents learning MM stochastic multi-armed bandits.
result Proved near-optimal behavior of proposed algorithms for group regret.