In this paper we investigate the Follow the Regularized Leader dynamics in sequential imperfect information games (IIG). We generalize existing results of Poincaré recurrence from normal-form games to zero-sum two-player imperfect information games and other sequential game settings. We then investigate how adapting th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Policy gradient method proves convergence in imperfect-information games.
JPS improves joint policies for multi-agent collaboration in imperfect information games.
Study learns optimal strategies in imperfect information games with self-play.
Algorithm learns NE in imperfect information games with imperfect feedback.
Paper solves learning imperfect-information games with fewer episodes.
DREAM learns optimal strategies in imperfect games without needing a simulator.
Study on liquidity and market efficiency in auction games with imperfect information.
RLCFR improves CFR's generalization in imperfect information games.
We introduce a new virtual environment for simulating a card game known as "Big 2". This is a four-player game of imperfect information with a relatively complicated action space (being allowed to play 1,2,3,4 or 5 card combinations from an initial starting hand of 13 cards). As such it poses a challenge for many curre…
Adaptive OMD reduces variance in learning optimal strategies for imperfect information games.
Counterfactual regret minimization (CFR) is the most popular algorithm on solving two-player zero-sum extensive games with imperfect information and achieves state-of-the-art performance in practice. However, the performance of CFR is not fully understood, since empirical results on the regret are much better than the …
We study pricing and superhedging strategies for game options in an imperfect market with default. We extend the results obtained by Kifer in \cite{Kifer} in the case of a perfect market model to the case of an imperfect market with default, when the imperfections are taken into account via the nonlinearity of the weal…
Cross-dimensional neural networks improve AI in Catan game.
First sample-efficient algorithm for learning EFCE in bandit feedback settings.
Investigates a Kyle model with imperfect information and risk aversion.
This paper tackles sample-efficient reinforcement learning for partially observable Markov games.
On April 13th, 2019, OpenAI Five became the first AI system to defeat the world champions at an esports game. The game of Dota 2 presents novel challenges for AI systems such as long time horizons, imperfect information, and complex, continuous state-action spaces, all challenges which will become increasingly central …
Bayesian probability theory is one of the most successful frameworks to model reasoning under uncertainty. Its defining property is the interpretation of probabilities as degrees of belief in propositions about the state of the world relative to an inquiring subject. This essay examines the notion of subjectivity by dr…
Partial-monitoring games constitute a mathematical framework for sequential decision making problems with imperfect feedback: The learner repeatedly chooses an action, opponent responds with an outcome, and then the learner suffers a loss and receives a feedback signal, both of which are fixed functions of the action a…
Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function representing discounted return. In this paper, we examine the role of these policy gradi…
From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made dramatic advances with artificial agents reaching superhuman performance in challenge domains like Go, Atari, and some variants of poker. A…
Real life hedging in the Black-Scholes model must be imperfect and if the stock's drift is higher than the risk free rate, leads to a profit on average. Hence the option price is examined as a fair game agreement between the parties, based on expected payoffs and a simple measure of risk. The resulting prices result in…
This work presents a game-theoretic method for AVs that handles imperfect communication and individual rewards.
Bayesian framework calibrates imperfect models using physics-informed priors and Hamiltonian Monte Carlo.
The paper tackles targeted attacks on rank aggregation methods, proving the fixed point of adversarial game.
We derive asset pricing formula for markets with incomplete information and subjective views.
Paper introduces a method to learn physics between digital twins using imperfect models.
Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior is an approximately optimal solution to an unknown decision problem. These tech…
Policy gradient and actor-critic algorithms form the basis of many commonly used training techniques in deep reinforcement learning. Using these algorithms in multiagent environments poses problems such as nonstationarity and instability. In this paper, we first demonstrate that standard softmax-based policy gradient c…
We present a novel methodology for predicting future outcomes that uses small numbers of individuals participating in an imperfect information market. By determining their risk attitudes and performing a nonlinear aggregation of their predictions, we are able to assess the probability of the future outcome of an uncert…
Solves a game between brokers and informed traders using stochastic differential equations.
When the available statistical information is imperfect, it is dangerous to follow standard optimisation procedures to construct an optimal portfolio, which usually leads to a strong concentration of the weights on very few assets. We propose a new way, based on generalised entropies, to ensure a minimal degree of dive…
In 2015, Google's DeepMind announced an advancement in creating an autonomous agent based on deep reinforcement learning (DRL) that could beat a professional player in a series of 49 Atari games. However, the current manifestation of DRL is still immature, and has significant drawbacks. One of DRL's imperfections is it…
Selective planning with imperfect models reduces harmful effects of model inadequacy.
Investors' strategic trading affects asset prices, modeled as a game.
In this paper the theory of semi-bounded rationality is proposed as an extension of the theory of bounded rationality. In particular, it is proposed that a decision making process involves two components and these are the correlation machine, which estimates missing values, and the causal machine, which relates the cau…
In this paper, we investigate Dimensionality reduction (DR) maps in an information retrieval setting from a quantitative topology point of view. In particular, we show that no DR maps can achieve perfect precision and perfect recall simultaneously. Thus a continuous DR map must have imperfect precision. We further prov…
Study proves value of non-Markovian games with partial, asymmetric info.
DefogGAN predicts hidden RTS game information to aid strategic decision-making.
Most approaches aiming to ensure a model's fairness with respect to a protected attribute (such as gender or race) assume to know the true value of the attribute for every data point. In this paper, we ask to what extent fairness interventions can be effective even when only imperfect information about the protected at…
Paper defines saddle points in asymmetric Dynkin games using martingale theory.
Financial markets, with their vast range of different investment opportunities, can be seen as a system of many different simultaneous games with diverse and often unknown levels of risk and reward. We introduce generalizations to the classic Kelly investment game [Kelly (1956)] that incorporates these features, and us…
The paper tackles learning from imperfect human feedback, especially in dueling bandit problems.
Study best-response learning dynamics in zero-sum polymatrix games under full and minimal information settings.
Imitation learning (IL) aims to learn an optimal policy from demonstrations. However, such demonstrations are often imperfect since collecting optimal ones is costly. To effectively learn from imperfect demonstrations, we propose a novel approach that utilizes confidence scores, which describe the quality of demonstrat…
We study analytically and numerically Minority Games in which agents may invest in different assets (or markets), considering both the canonical and the grand-canonical versions. We find that the likelihood of agents trading in a given asset depends on the relative amount of information available in that market. More s…
New framework uses tempered optimism to handle imperfect experts in online learning.