New algorithms optimize decision rules in strategic scenarios, minimizing prediction risk and incentivizing better outcomes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In principal-agent models, a principal offers a contract to an agent to perform a certain task. The agent exerts a level of effort that maximizes her utility. The principal is oblivious to the agent's chosen level of effort, and conditions her wage only on possible outcomes. In this work, we consider a model in which t…
Equilibrium found for multi-agent trading with transaction costs.
New algorithm for maximizing revenue in multinomial logistic bandits.
We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on the average vectorial outcome. The problem models applications such as multi-objective optimization, maximum entropy exploration, and constrai…
Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior is an approximately optimal solution to an unknown decision problem. These tech…
Social learning can make financial markets inefficient, but individual learning can fix this.
Explains agent behavior through intended outcomes in reinforcement learning.
The game theory techniques are used to find the equilibrium of a market. Game theory refers to the ways in which strategic interactions among economic agents produce outcomes with respect to the preferences (or utilities) of those agents, where the outcomes in question might have been intended by none of the agents. Th…
Algorithms are often used to produce decision-making rules that classify or evaluate individuals. When these individuals have incentives to be classified a certain way, they may behave strategically to influence their outcomes. We develop a model for how strategic agents can invest effort in order to change the outcome…
Agents learn to outperform in trading by using past and current prices.
InfoTree improves reinforcement learning by optimizing tool use with a greedy submodular approach.
We consider the problem of belief aggregation: given a group of individual agents with probabilistic beliefs over a set of uncertain events, formulate a sensible consensus or aggregate probability distribution over these events. Researchers have proposed many aggregation methods, although on the question of which is be…
We tackle the blackbox issue of deep neural networks in the settings of reinforcement learning (RL) where neural agents learn towards maximizing reward gains in an uncontrollable way. Such learning approach is risky when the interacting environment includes an expanse of state space because it is then almost impossible…
Study on learning strategies in matching markets with uncertain preferences.
A method to train multi-agent reinforcement learning models without intrinsic rewards.
This paper proposes a formal approach to online learning and planning for agents operating in a priori unknown, time-varying environments. The proposed method computes the maximally likely model of the environment, given the observations about the environment made by an agent earlier in the system run and assuming know…
MA-COPP predicts multi-agent system outcomes using data from a different policy, with probabilistic guarantees.
Reinforcement learning refers to a group of methods from artificial intelligence where an agent performs learning through trial and error. It differs from supervised learning, since reinforcement learning requires no explicit labels; instead, the agent interacts continuously with its environment. That is, the agent sta…
There is tremendous interest in precision medicine as a means to improve patient outcomes by tailoring treatment to individual characteristics. An individualized treatment rule formalizes precision medicine as a map from patient information to a recommended treatment. A treatment rule is defined to be optimal if it max…
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
Optimizes long-term social welfare in recommender systems by matching users to providers.
DEBIAS learns causal effects from psychiatric longitudinal data by optimizing outcome weights.
Unified framework for human-like decision making in various sequential tasks.
ISMCTS-BR learns best responses in large games, approximating worst-case performance.
A new framework for causal inference in networked settings.
Active inference minimizes expected free energy for optimal behavior.
Optimizes non-linear outcomes from summed contributions.
We propose a probabilistic modeling framework for learning the dynamic patterns in the collective behaviors of social agents and developing profiles for different behavioral groups, using data collected from multiple information sources. The proposed model is based on a hierarchical Bayesian process, in which each obse…
Framework trains safe agents avoiding deceptive behavior.
The paper shows how shared random seeds can reduce variance in machine learning evaluations.
A new method for optimizing stakes in a single event with multiple outcomes.
Optimizes treatment duration to maximize quality-adjusted lifetime.
Efficient sequential matching of supply and demand is a problem of interest in many online to offline services. For instance, Uber, Lyft, Grab for matching taxis to customers; Ubereats, Deliveroo, FoodPanda etc for matching restaurants to customers. In these online to offline service problems, individuals who are respo…
Study shows how high-budget agents can manipulate prediction markets.
We discuss here the mean-field theory for a cellular automata model of meta-learning. The meta-learning is the process of combining outcomes of individual learning procedures in order to determine the final decision with higher accuracy than any single learning method. Our method is constructed from an ensemble of inte…
The paper analyzes optimal dealer strategies in agent-based market models.
New algorithm tackles unknown utility network resource allocation.
Develops an LLM-based agent for superior cryptocurrency trading.
We consider a framework involving behavioral economics and machine learning. Rationally inattentive Bayesian agents make decisions based on their posterior distribution, utility function and information acquisition cost Renyi divergence which generalizes Shannon mutual information). By observing these decisions, how ca…
The autonomous trading agent is one of the most actively studied areas of artificial intelligence to solve the capital market portfolio management problem. The two primary goals of the portfolio management problem are maximizing profit and restrainting risk. However, most approaches to this problem solely take account …
We are working to develop automated intelligent agents, which can act and react as learning machines with minimal human intervention. To accomplish this, an intelligent agent is viewed as a question-asking machine, which is designed by coupling the processes of inference and inquiry to form a model-based learning unit.…
This paper presents a computational model for conceptual shifts, based on a novelty metric applied to a vector representation generated through deep learning. This model is integrated into a co-creative design system, which enables a partnership between an AI agent and a human designer interacting through a sketching c…
Study on blackjack reinforcement learning performance with varying deck sizes.
Algorithm maximizes total reward in multi-agent bandits with adversarial corruptions.
The paper targets optimal interventions for long-term outcomes using imputed data and policy learning.
We provide an axiomatic foundation for the representation of numéraire-invariant preferences of economic agents acting in a financial market. In a static environment, the simple axioms turn out to be equivalent to the following choice rule: the agent prefers one outcome over another if and only if the expected (under t…
RL agent learns to avoid market spoofing.