GPT learns a causal world model from token predictions, validated in game sequences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Zero-sum games such as chess and poker are, abstractly, functions that evaluate pairs of agents, for example labeling them `winner' and `loser'. If the game is approximately transitive, then self-play generates sequences of agents of increasing strength. However, nontransitive games, such as rock-paper-scissors, can ex…
We study a variant of the source identification game with training data in which part of the training data is corrupted by an attacker. In the addressed scenario, the defender aims at deciding whether a test sequence has been drawn according to a discrete memoryless source , whose statistics are known to hi…
In this paper we consider Dynkin's games with payoffs which are functions of an underlying process. Assuming extended weak convergence of underlying processes to a limit process we prove convergence Dynkin's games values corresponding to to the Dynkin's game…
We study the convergence of Nash equilibria in a game of optimal stopping. If the associated mean field game has a unique equilibrium, any sequence of -player equilibria converges to it as . However, both the finite and infinite player versions of the game often admit multiple equilibria. We show that me…
Improved FTPL algorithm reduces regret in predictable minimax games.
Study shows aperiodic sequences enhance Parrondo's effect, with Thue-Morse outperforming others.
Q*BERT learns to navigate text-based games by building a knowledge graph.
Low precision networks in the reinforcement learning (RL) setting are relatively unexplored because of the limitations of binary activations for function approximation. Here, in the discrete action ATARI domain, we demonstrate, for the first time, that low precision policy distillation from a high precision network pro…
Algorithm learns Nash equilibria in stochastic games using entropy-regularized policies.
Quantum strategy optimizes wealth growth in a double-or-nothing game.
This paper introduces a new class of Dynkin games, where the two players are allowed to make their stopping decisions at a sequence of exogenous Poisson arrival times. The value function and the associated optimal stopping strategy are characterized by the solution of a backward stochastic differential equation. The pa…
We present a modification of the so-called Parrondo's paradox where one is allowed to choose in each turn the game that a large number of individuals play. It turns out that, by choosing the game which gives the highest average earnings at each step, one ends up with systematic loses, whereas a periodic or random seque…
The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such as external regret do not take into account. We revisit the notion of policy regret and first show that there are online learning settings …
Algorithm learns to play against unknown opponents in sequential games.
Algorithm minimizes regret and converges to equilibria in Markov games.
A game theory study on optimal hiding and searching strategies in discrete locations.
This work finds mixed equilibria in zero-sum games using interacting particle dynamics.
Interpretability has arisen as a key desideratum of machine learning models alongside performance. Approaches so far have been primarily concerned with fixed dimensional inputs emphasizing feature relevance or selection. In contrast, we focus on temporal modeling and the problem of tailoring the predictor, functionally…
Paper compares two forecasters using novel online inference methods.
Improved upper bound for online calibrated forecasting of binary sequences.
Study on learning strategies in adaptive Markov games with policy regret as metric.
Existing imitation learning approaches often require that the complete demonstration data, including sequences of actions and states, are available. In this paper, we consider a more realistic and difficult scenario where a reinforcement learning agent only has access to the state sequences of an expert, while the expe…
Adaptive OMD reduces variance in learning optimal strategies for imperfect information games.
Generating music medleys is about finding an optimal permutation of a given set of music clips. Toward this goal, we propose a self-supervised learning task, called the music puzzle game, to train neural network models to learn the sequential patterns in music. In essence, such a game requires machines to correctly sor…
Recent literature on online learning has focused on developing adaptive algorithms that take advantage of a regularity of the sequence of observations, yet retain worst-case performance guarantees. A complementary direction is to develop prediction methods that perform well against complex benchmarks. In this paper, we…
MAXMINLCB optimizes unknown target functions with preference feedback using a Stackelberg game approach.
Activities in reinforcement learning (RL) revolve around learning the Markov decision process (MDP) model, in particular, the following parameters: state values, V; state-action values, Q; and policy, pi. These parameters are commonly implemented as an array. Scaling up the problem means scaling up the size of the arra…
We study a risk sensitive control version of the lifetime ruin probability problem. We consider a sequence of investments problems in Black-Scholes market that includes a risky asset and a riskless asset. We present a differential game that governs the limit behavior. We solve it explicitly and use it in order to find …
We show that prices and shortfall risks of game (Israeli) barrier options in a sequence of binomial approximations of the Black--Scholes (BS) market converge to the corresponding quantities for similar game barrier options in the BS market with path dependent payoffs and the speed of convergence is estimated, as well. …
New discrete-time model shows insider trading dynamics.
New framework connects online learning to statistical learning for better generalization bounds.
OpenAlpha validates decentralized capital strategies using game theory and market aggregation.
New framework ensures valid uncertainty estimates for any data stream changes.
In order to communicate, humans flatten a complex representation of ideas and their attributes into a single word or a sentence. We investigate the impact of representation learning in artificial agents by developing graph referential games. We empirically show that agents parametrized by graph neural networks develop …
We propose a principled method for kernel learning, which relies on a Fourier-analytic characterization of translation-invariant or rotation-invariant kernels. Our method produces a sequence of feature maps, iteratively refining the SVM margin. We provide rigorous guarantees for optimality and generalization, interpret…
Sequence prediction models can be learned from example sequences with a variety of training algorithms. Maximum likelihood learning is simple and efficient, yet can suffer from compounding error at test time. Reinforcement learning such as policy gradient addresses the issue but can have prohibitively poor exploration …
Paper defines a new dimension to measure self-directed learning complexity.
Generative adversarial networks (GANs) are powerful tools for learning generative models. In practice, the training may suffer from lack of convergence. GANs are commonly viewed as a two-player zero-sum game between two neural networks. Here, we leverage this game theoretic view to study the convergence behavior of the…
New method constructs confidence sets for GLMs via game theory.
Study shows how adaptive market agents can lead to persistent overpricing in financial markets.
We consider the setting of online linear regression for arbitrary deterministic sequences, with the square loss. We are interested in the aim set by Bartlett et al. (2015): obtain regret bounds that hold uniformly over all competitor vectors. When the feature sequence is known at the beginning of the game, they provide…
We provide a new approach to training neural models to exhibit transparency in a well-defined, functional manner. Our approach naturally operates over structured data and tailors the predictor, functionally, towards a chosen family of (local) witnesses. The estimation problem is setup as a co-operative game between an …
We argue that the existing regret matchings for Nash equilibrium approximation conduct "jumpy" strategy updating when the probabilities of future plays are set to be proportional to positive regret measures. We propose a geometrical regret matching which features "smooth" strategy updating. Our approach is simple, intu…
Paper proposes GANs for generating business process suffixes and remaining times.
The firefighter game problem on locally finite connected graphs was introduced by Bert Hartnell. The game on a graph can be described as follows: let be a sequence of positive integers; an initial fire starts at a finite set of vertices; at each (integer) time , vertices which are not on fire b…
We address the issue of limit cycling behavior in training Generative Adversarial Networks and propose the use of Optimistic Mirror Decent (OMD) for training Wasserstein GANs. Recent theoretical results have shown that optimistic mirror decent (OMD) can enjoy faster regret rates in the context of zero-sum games. WGANs …
In many predictive decision-making scenarios, such as credit scoring and academic testing, a decision-maker must construct a model that accounts for agents' propensity to "game" the decision rule by changing their features so as to receive better decisions. Whereas the strategic classification literature has previously…