Paper tackles hidden game problem in AI alignment and language games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Gradient Descent Ascent converges to von-Neumann solution in hidden zero-sum games.
DefogGAN predicts hidden RTS game information to aid strategic decision-making.
The paper explains how simple methods can converge to optimal solutions in complex neural games.
Contributions: Prior studies on education have mostly followed the model of the cross sectional study, namely, examining the pretest and the posttest scores. This paper shows that students' knowledge throughout the intervention can be estimated by time series analysis using a hidden Markov model. Background: Analyzing …
Study explores reinforcement learning in a complex game environment, analyzing rule inference and policy learning.
Recent breakthroughs in AI for multi-agent games like Go, Poker, and Dota, have seen great strides in recent years. Yet none of these games address the real-life challenge of cooperation in the presence of unknown and uncertain teammates. This challenge is a key game mechanism in hidden role games. Here we develop the …
CausalGame benchmarks LLM agents' causal thinking in games.
Market makers optimize bid/ask quotes under hidden Markov chain uncertainty.
Algorithm finds ε-equilibrium policies for multi-agent Markov games with hidden low-rank structure.
Individuals, or organizations, cooperate with or compete against one another in a wide range of practical situations. Such strategic interactions are often modeled as games played on networks, where an individual's payoff depends not only on her action but also on that of her neighbors. The current literature has large…
Activities in reinforcement learning (RL) revolve around learning the Markov decision process (MDP) model, in particular, the following parameters: state values, V; state-action values, Q; and policy, pi. These parameters are commonly implemented as an array. Scaling up the problem means scaling up the size of the arra…
Deep Q-learning is investigated as an end-to-end solution to estimate the optimal strategies for acting on time series input. Experiments are conducted on two idealized trading games. 1) Univariate: the only input is a wave-like price time series, and 2) Bivariate: the input includes a random stepwise price time series…
Study on multi-agent decision making complexity, showing sample efficiency gaps.
A game theory study on optimal hiding and searching strategies in discrete locations.
We prove a general connection between the communication complexity of two-player games and the sample complexity of their multi-player locally private analogues. We use this connection to prove sample complexity lower bounds for locally differentially private protocols as straightforward corollaries of results from com…
This paper presents OptNet, a network architecture that integrates optimization problems (here, specifically in the form of quadratic programs) as individual layers in larger end-to-end trainable deep networks. These layers encode constraints and complex dependencies between the hidden states that traditional convoluti…
We propose dynamical systems trees (DSTs) as a flexible class of models for describing multiple processes that interact via a hierarchy of aggregating parent chains. DSTs extend Kalman filters, hidden Markov models and nonlinear dynamical systems to an interactive group scenario. Various individual processes interact a…
New architecture separates object state and behavior for better game dynamics.
IGGP learns game rules from varying quality game play, finding no overall trend.
In this paper we study the nonzero-sum Dynkin game in continuous time which is a two player non-cooperative game on stopping times. We show that it has a Nash equilibrium point for general stochastic processes. As an application, we consider the problem of pricing American game contingent claims by the utility maximiza…
Do you remember your first video game console? We remember ours. Decades ago, they provided hours of entertainment. Now, we have repurposed them to solve dynamic and stochastic optimization problems. With deep reinforcement learning methods posting superhuman performance on a wide range of Atari games, we consider the …
Develops a deep metric learning approach for detecting bugs in video games.
Wasserstein GANs are shown to have hidden convexity, enabling exact solutions with convex optimization.
This paper addresses the problem of multi-agent inverse reinforcement learning (MIRL) in a two-player general-sum stochastic game framework. Five variants of MIRL are considered: uCS-MIRL, advE-MIRL, cooE-MIRL, uCE-MIRL, and uNE-MIRL, each distinguished by its solution concept. Problem uCS-MIRL is a cooperative game in…
In this paper we propose and analyze a class of -player stochastic games that include finite fuel stochastic games as a special case. We first derive sufficient conditions for the Nash equilibrium (NE) in the form of a verification theorem. The associated Quasi-Variational-Inequalities include an essential game comp…
Study optimal investment and consumption strategies for competitive agents with habit formation.
We introduce a deep, generative autoencoder capable of learning hierarchies of distributed representations from data. Successive deep stochastic hidden layers are equipped with autoregressive connections, which enable the model to be sampled from quickly and exactly via ancestral sampling. We derive an efficient approx…
New measure of feature influence in classification problems considering feature dependencies.
The paper tackles Nash-regret minimization in congestion games with bandit feedback.
Study of portfolio management under relative performance concerns using mean field games.
Solves specific mean-field game equations with ODEs.
Potential games, originally introduced in the early 1990's by Lloyd Shapley, the 2012 Nobel Laureate in Economics, and his colleague Dov Monderer, are a very important class of models in game theory. They have special properties such as the existence of Nash equilibria in pure strategies. This note introduces graphical…
Federated learning linked to mean-field games for large-scale learning.
We consider the problem of finding stationary Nash equilibria (NE) in a finite discounted general-sum stochastic game. We first generalize a non-linear optimization problem from Filar and Vrieze [2004] to a -player setting and break down this problem into simpler sub-problems that ensure there is no Bellman error fo…
Algorithm learns from changing zero-sum games with no regret.
Linear recurrent networks explain reinforcement learning performance in partially observable settings.
Recent successes of game-theoretic formulations in ML have caused a resurgence of research interest in differentiable games. Overwhelmingly, that research focuses on methods and upper bounds on their speed of convergence. In this work, we approach the question of fundamental iteration complexity by providing lower boun…
In this paper, we study the problem of learning the set of pure strategy Nash equilibria and the exact structure of a continuous-action graphical game with quadratic payoffs by observing a small set of perturbed equilibria. A continuous-action graphical game can possibly have an uncountable set of Nash euqilibria. We p…
The paper solves investment problems with uncertain factors using game theory.
New research shows no-regret learning is impossible in Markov games under certain assumptions.
A quantum financial approach to finite games of strategy is addressed, with an extension of Nash's theorem to the quantum financial setting, allowing for an entanglement of games of strategy with two-period financial allocation problems that are expressed in terms of: the consumption plans' optimization problem in pure…
Unified framework for estimating reward functions in competitive games.
Deep neural network solves large multi-agent games for Markovian Nash equilibrium.
Study equilibrium consumption habits in a large population using mean field games.
We introduce CSE for MLSF games and devise online learning algorithms for achieving no-external Stackelberg-regret.
Deep Reinforcement Learning automates match-3 game testing.
A new algorithm learns policies from batch data in hierarchical RL.