Transformers learn to play games in-context, proving Nash equilibrium.
problem Understanding in-context game-playing capabilities of pre-trained transformers.
method Theoretical guarantees and constructional results for transformer architecture in multi-agent games.
result Pre-trained transformers can learn Nash equilibrium in-context for two-player zero-sum games.
Paper uses DRL to evaluate NHL players based on game context.
problem Limited ability to model player performance in game context.
method Applied Deep Reinforcement Learning to learn Q-function from NHL play-by-play data.
result Game Impact Metric (GIM) correlates with player success and future salary.
New games model strategic interactions in incomplete information settings.
problem Modeling strategic interactions in incomplete information settings.
method Introduced new games that map input to private player types, aggregate strategies, and converge to near-Nash equilibria.
result Games can recover meaningful strategic interactions from real data.
Improved RL for TBGs by pruning irrelevant tokens and bootstrapping.
problem RL methods fail to generalize in TBGs with small data.
method CREST for irrelevant token removal, bootstrapped Q-learning.
result Improved generalization in unseen TextWorld games.
This work trains a model to generate all valid commands for text-based games.
problem Generating valid commands for text-based games.
method Training generative models on a dataset of text-based game contexts.
result The best model can generate valid commands unseen at training and achieves high F1 score.
Bayesian optimization finds game equilibria efficiently.
problem Finding game equilibria in derivative-free, expensive contexts.
method Bayesian optimization framework with sequential sampling decisions.
result Equilibria can be found reliably at a lower cost.
Develops a deep metric learning approach for detecting bugs in video games.
problem Automated detection of bugs in video games.
method State-State Siamese Networks (S3N) for deep metric learning.
result S3N learns meaningful embeddings to identify various types of bugs.
Educational game on crypto investment helps students grasp macroeconomics.
problem Weak connections between microeconomic decision-making and macroeconomic concepts in classroom games.
method Design and study of an educational game on cryptocurrency investment.
result Engages students in understanding macroeconomics through incentivized individual investment decisions.
Paper predicts player retention in mobile games quickly.
problem Predicting player retention in free-to-play mobile games.
method Heuristic modeling to build simple retention prediction rules.
result Heuristic approach achieves comparable performance to common algorithms.
Study on mean field games with singular controls and their applications.
problem Optimal productivity expansion in dynamic oligopolies.
method Existence and uniqueness of mean field equilibria through nonlinear equations, Abelian limit for discounted and ergodic games.
result Valid connection between discounted and ergodic games, approximation of Nash equilibria.
Gradient methods solve Stackelberg games in high-dimensional, continuous decision spaces.
problem Solving Stackelberg games with continuous, high-dimensional decisions.
method Gradient methods for searching numerical solutions.
result Time and space scalability of gradient methods discussed.
Introduces Space Fortress to test RL algorithms' context and time sensitivity.
problem RL benchmarks lack context-dependent shifts and temporal sensitivity.
method Introduces Space Fortress as a new RL benchmark.
result Existing RL algorithms fail on Space Fortress due to context insensitivity and reward sparsity.
We develop an option pricing model based on a tug-of-war game. This two-player zero-sum stochastic differential game is formulated in the context of a multi-dimensional financial market. The issuer and the holder try to manipulate asset price processes in order to minimize and maximize the expected discounted reward. W…
Study on hypothesis testing games with adversarial classification, showing convergence rates.
problem Adversarial classification in hypothesis testing.
method Mixed strategy Nash equilibria analysis, concentration phenomena examination.
result Exponential rates of convergence of classification errors at equilibrium.
This work compresses reinforcement learning models for Atari games, improving localization.
problem Expensive deep neural networks in reinforcement learning.
method Model compression, global max-pooling, Actor-Mimic, weakly supervised localization.
result Compression reduces model size to 3% of original, enabling object localization.
New approach for making machine learning models understandable over time.
problem Making machine learning models interpretable for temporal data.
method Cooperative game between predictor and explainer for temporal models.
result Predictor can be tailored to fit interpretable temporal models.
A system estimates delayed context for online scoring using convex optimization.
problem Estimating agent scores with delayed context information.
method Online convex game between agent and system; leveraging correlation function.
result Error in score estimate is small if online convex game has low regret.
Study N N N -player and mean-field games in Itô-diffusion markets with competitive or homophilous interactions.
problem Optimal portfolio choice in a common market with N N N interacting players. method Analyzes N N N -player and mean-field games in incomplete and complete markets with CARA utilities and random risk tolerances. result Derives explicit or closed-form solutions for equilibrium processes and game values.
The paper explains the equity premium without probabilistic assumptions.
problem Understanding the equity premium and CAPM without probabilistic assumptions.
method Develops game-theoretic probability in continuous-time financial markets.
result Derives a simple expression for the equity premium and a version of CAPM.
A new algorithm calculates optimal strategies for two-player zero-sum games.
problem Computing the optimal strategies for two-player zero-sum games.
method Extending successive relaxation to two-player zero-sum games and developing a generalized minimax Q-learning algorithm.
result The proposed algorithm converges and effectively computes optimal strategies.
Pref-SHAP explains preferences using Shapley values.
problem Challenging problem of preference explanation in machine learning.
method Shapley value-based model explanation framework for pairwise comparison data.
result Richer and more insightful explanations obtained over baseline.
New measure of feature influence in classification problems considering feature dependencies.
problem Measuring the influence of features in classification problems with dependencies.
method Developed a new measure based on cooperative game theory, providing axiomatic characterization and demonstrating its equivalence to the Banzhaf-Owen value.
result The proposed influence measure effectively characterizes feature importance in classification problems with feature dependencies.
New algorithms achieve logarithmic regret in KL-regularized Markov games.
problem Improving sample efficiency in game-theoretic settings with KL regularization.
method Developed OMG and SOMG algorithms for matrix and Markov games, using best response sampling and superoptimistic bonuses.
result Logarithmic regret in T T T that scales inversely with KL regularization strength β β β . Model predicts user engagement and survival time in games.
problem Building data-efficient engagement models for diverse games.
method Data-driven approach using minimal metrics.
result Joint estimates of survival time and churn probability.
The paper extends macroscopic market making to stochastic games, revealing properties and solving equations.
problem Price competition among market makers in a stochastic game setting.
method Extension of macroscopic market making framework to stochastic games, introducing multidimensional characteristic equations.
result New well-posedness results for forward-backward stochastic differential equations.
Paper tackles inverse reinforcement learning with non-optimal demonstrations in zero-sum games.
problem Inverse reinforcement learning with sub-optimal expert demonstrations in zero-sum games.
method Introduces a new objective function and algorithm to find reward function and strategies without decoupling agents.
result Demonstrates recovery of reward functions and strategies with good quality from sub-optimal expert demonstrations.
We study ranking quantilized mean-field games to select top-performing agents.
problem Selecting top-performing agents in competitive scenarios.
method Developed two formulations: target-based and threshold-based, and provided analytic and semi-explicit solutions.
result Analytic and semi-explicit solutions for quantilized mean-field consistency conditions.
EigenGame improves eigendecomposition by offering unbiased updates for larger datasets.
problem Minibatch bias in EigenGame limits convergence and parallelism.
method Proposed unbiased stochastic update for EigenGame.
result Asymptotic equivalence to EigenGame, greater parallelism, and improved performance.
Study proposes new OPE estimators for two-player zero-sum games.
problem Evaluating new policies using historical data from a different policy in multi-player zero-sum games.
method Doubly robust and double reinforcement learning estimators to project exploitability.
result Prove exploitability estimation error bounds and regret bounds for policy profiles.
Two strategic agents track their portfolios, influencing each other's trading targets.
problem Strategic competition in portfolio tracking with price impact.
method Stochastic linear quadratic differential game with terminal state constraints.
result Unique open-loop Nash equilibrium strategies emerge based on price impact types.
Paper studies zero-sum games with noisy observations and identifies equilibrium conditions.
problem Zero-sum games with noisy observations of the leader's actions.
method Analyzes the equilibrium of games with noisy action observability, identifies necessary conditions for uniqueness, and investigates the cardinality of best responses.
result The noisy observations significantly impact the cardinality of the follower's set of best responses, and under certain conditions, this set becomes a singleton almost surely.
Study of pursuit-evasion game on sphere and its relation to planar Apollonius circle.
problem Analyzing pursuit-evasion game on a sphere and its properties.
method Extending classical planar pursuit-evasion game to spherical geometry, studying equilibrium intercept points and their relation to Apollonius domain.
result Condition for intercept point to belong to Apollonius domain on sphere, analogous to planar game.
As we show using the notion of equilibrium in the theory of infinite sequential games, bubbles and escalations are rational for economic and environmental agents, who believe in an infinite world. This goes against a vision of a self regulating, wise and pacific economy in equilibrium. In other words, in this context, …
This paper examines transitions in sniping behavior among algorithmic traders, finding new profitable strategies.
problem Understanding transitions from sure to probabilistic sniping in competitive algorithmic trading environments.
method Reinterpretation and extension of Menkveld and Zoican's stylized game, analysis of repeated games, sequential statistical testing.
result Probabilistic sniping can be profitable in certain conditions, resembling the prisoner's dilemma.
Competition has been introduced in the electricity markets with the goal of reducing prices and improving efficiency. The basic idea which stays behind this choice is that, in competitive markets, a greater quantity of the good is exchanged at a lower and a lower price, leading to higher market efficiency. Electricity …
This study compares global vs local observation and action representations for DRL in RTS games.
problem Improving Deep Reinforcement Learning performance in RTS games.
method Comparing two observation and action representations in μRTS.
result Local representation outperforms global representation in resource harvesting tasks.
Unified approach to fair online learning with stochastic contexts.
problem Fairness in online learning with unknown sensitive contexts.
method Adapting Blackwell's approachability theory to handle unknown contexts' distributions.
result Characterization of optimal trade-off between fairness and performance objectives.
OMD improves training GANs by avoiding limit cycling.
problem Limit cycling behavior in training GANs.
method Optimistic Mirror Decent (OMD) for training Wasserstein GANs.
result OMD addresses limit cycling in WGANs and converges to an equilibrium.
Study shows how multiple traders can trade together without excessive price impact.
problem Coordination issues in trading to exploit a common signal.
method Closed-loop Nash competition model for stochastic differential games.
result Excessive trading reduced but not significantly for practical parameters.
Broadens Jourdain and Martini's method to non-linear stochastic processes.
problem Applying pricing methods to non-linear stochastic processes.
method Analyzes from probabilistic and analytic viewpoints, extending Jourdain and Martini's method.
result Broadens applicability of pricing methods to non-linear frameworks.
Paper introduces new metrics for evaluating player actions in soccer.
problem Traditional metrics focus on rare actions and ignore contextual information.
method New language and framework for valuing player actions considering context.
result Total offensive and defensive contributions can be quantified.
New method uses game theory to rate generative models.
problem Evaluating generative models' performance.
method Tournaments between generators and discriminators to summarize outcomes.
result Tournament win rate and skill rating provide effective model evaluations.
The study explores when parametric models enhance reinforcement learning, validating a hypothesis on Atari games.
problem When and how to use parametric models in reinforcement learning.
method Comparison of parametric models and experience replay, validating a hypothesis on Atari games.
result Replay-based algorithms can be competitive or superior to model-based algorithms under suitable conditions.
Research explores how interconnected systems synchronize and how to control their behavior.
problem Understanding and controlling the behavior of interconnected dynamical systems.
method Mean field games approach applied to controlled coupled oscillators.
result Developed methods to predict and influence emergent phenomena in interconnected systems.
New approach for neural models to be transparent over structured data.
problem Training neural models to be transparent in a functional manner.
method Setup as a cooperative game between a predictor and witnesses, encouraging local agreement.
result The predictor remains globally powerful while agreeing locally with witnesses.
Study optimal investment-reinsurance strategies in equity-linked insurance products using Stackelberg game theory.
problem Optimizing investment and reinsurance strategies in equity-linked insurance products with capital guarantees.
method Modelled as a Stackelberg game where reinsurer acts as leader and insurer as follower, with general utility functions and power utility functions analyzed.
result Derive Stackelberg equilibrium for general utility functions and calculate it explicitly for power utility functions, finding reinsurer optimizes premium to incentivize maximal reinsurance purchase.
New framework compares two stochastic learning dynamics in games.
problem Inability to distinguish between different learning rules leading to the same steady-state behavior.
method Developed a framework for comparative analysis of stochastic learning dynamics with different update rules.
result Identified distinct behaviors in the paths to stochastically stable states for LLL and ML.
DORIS algorithm achieves no-regret learning in Markov games with adversarial opponents.
problem Decentralized policy learning in Markov games with nonstationary opponents.
method DORIS algorithm using optimistic hyperpolicy mirror descent.
result Achieves K \sqrt{K} K -regret in general function approximation.