This paper shows hedging algorithms improve performance in repeated matrix games.
problem Improving multi-agent learning algorithms in repeated matrix games.
method Develops and experiments with hedging algorithms combining a top-level and a set of basic algorithms.
result Well-selected hedging algorithms outperform previous MAL algorithms on repeated matrix games.
Algorithm finds near-optimal strategy in changing zero-sum games.
problem Finding near-optimal strategy in changing zero-sum games.
method Designing an algorithm with small NE regret for online matrix games.
result Achieves near-optimal dependence on the number of rounds and number of actions.
Optimal strategies are found for a repeated betting game using diffusion approximation.
problem Finding optimal strategies for a repeated betting game with i.i.d. outcomes.
method Constructing a diffusion approximation of the repeated game and analyzing the wealth share process.
result Necessary and sufficient conditions for the wealth share process to be transient or recurrent are derived.
Study explores algorithmic collusion in repeated games using various learning dynamics.
problem Understanding algorithmic collusion in repeated games with different learning dynamics.
method Examines Q Q Q -learning, gradient learning, and other dynamics in a general repeated game setting. result Characterizes the set of payoff vectors achievable by these dynamics, revealing possibilities for collusion.
Algorithm learns from changing zero-sum games with no regret.
problem Learning in time-varying zero-sum games.
method Developed a single parameter-free algorithm with three performance measures.
result Algorithm recovers best known results for fixed games and adapts to non-stationarity.
LAFF algorithm balances adaptability and non-exploitability in repeated games.
problem Low regret in repeated games against unknown opponent classes.
method LAFF algorithm searches within sub-algorithms optimal for each opponent class and uses a punishment policy for exploitation.
result LAFF guarantees sublinear regret uniformly over possible opponents, except exploitative ones, for which it guarantees linear regret.
This paper examines transitions in sniping behavior among algorithmic traders, finding new profitable strategies.
problem Understanding transitions from sure to probabilistic sniping in competitive algorithmic trading environments.
method Reinterpretation and extension of Menkveld and Zoican's stylized game, analysis of repeated games, sequential statistical testing.
result Probabilistic sniping can be profitable in certain conditions, resembling the prisoner's dilemma.
New method detects heuristics in complex game strategies.
problem Understanding decision-making in games with infinite strategy spaces.
method Introducing decoupled strategies to detect convergence towards Nash equilibria.
result Predictive measure ΔD reveals participants' actions with high success rate.
Algorithm learns to play against unknown opponents in sequential games.
problem Designing strategies for a learner to interact with an unknown opponent in repeated sequential games.
method Kernel-based regularity assumptions and a novel algorithm combining bilevel optimization and online learning.
result Algorithm achieves sublinear regret guarantees and is effective in specific game settings.
R2-B2 optimizes game interactions with recursive reasoning.
problem Optimizing interactions between boundedly rational agents with unknown payoff functions.
method Recursive Reasoning-Based Bayesian Optimization (R2-B2) for repeated games.
result R2-B2 achieves faster asymptotic convergence to no regret than non-recursive methods.
New algorithms converge faster to Nash equilibrium in zero-sum games with bandit feedback.
problem Learning in zero-sum games with bandit feedback without communication.
method Developed two uncoupled algorithms achieving optimal rate of Ω ( T − 1 / 4 ) Ω(T^{-1/4}) Ω ( T − 1/4 ) . result Achieved optimal rate of Ω ( T − 1 / 4 ) Ω(T^{-1/4}) Ω ( T − 1/4 ) for convergence of policy profiles to Nash equilibrium. No-regret learning fails to converge to Nash equilibria in mixed strategies.
problem Limiting behavior of mixed strategies in repeated games.
method Study of optimal no-regret learning algorithms for 2x2 competitive games.
result Limiting mixed strategies cannot converge to Nash equilibria under mean-based and monotonic updates.
We consider the dynamics of player's strategies in repeated market games, where the selection of strategies is determined by a learning model. Prior theoretical analysis and experimental data show that after large number of plays the average number of agents who decide to enter, per round of the game, approaches the ma…
We describe an approximate dynamic programming (ADP) approach to compute approximations of the optimal strategies and of the minimal losses that can be guaranteed in discounted repeated games with vector-valued losses. Such games prominently arise in the analysis of regret in repeated decision-making in adversarial env…
A study on how a principal can incentivize an agent to make better decisions in a repeated game.
problem Optimizing a principal's utility in a misaligned principal-agent bandit game.
method Developed nearly optimal learning algorithms for the principal's regret in multi-armed and linear contextual settings.
result The principal can iteratively learn an incentive policy to maximize her total utility.
New algorithm learns from noisy and correlated game outcomes.
problem Learning to play a repeated multi-agent game with unknown reward function.
method GP-MW algorithm using Gaussian processes and multiplicative weight method.
result Novel kernel-dependent regret bounds comparable to full information settings.
We consider Blackwell approachability, a very powerful and geometric tool in game theory, used for example to design strategies of the uninformed player in repeated games with incomplete information. We extend this theory to "generalized quitting games" , a class of repeated stochastic games in which each player may ha…
We consider regret minimization in repeated games with non-convex loss functions. Minimizing the standard notion of regret is computationally intractable. Thus, we define a natural notion of regret which permits efficient optimization and generalizes offline guarantees for convergence to an approximate local optimum. W…
Algorithm improves RL model selection for repeated games with utility maximization.
problem Optimal policy learning in repeated games with unknown opponent strategy.
method Proposes MRBEAR for average reward RL, applying to utility maximization in repeated games.
result Regret bound shows linear dependence on number of model classes in average reward RL.
New algorithms improve on bandit feedback in matrix games with unknown payoff matrices.
problem Improving performance in matrix games with unknown payoff matrices and bandit feedback.
method Regret analyses of variants of UCB and K-learning.
result New algorithms achieve lower regret compared to adversarial bandit algorithms.
New algorithm improves game learning with randomised optimism.
problem Learning in matrix games with unknown payoffs and bandit feedback.
method Integrates evolutionary algorithms into bandit framework for randomised optimism.
result Achieves sublinear regret, outperforming classical methods.
Game theory helps machine learn better from adversarial queries.
problem Adversarial evasion in machine learning prediction.
method Repeated Bayesian Sequential Game to balance classifier selection and query type.
result Learner selects appropriate classifier for clean vs. adversarial queries.
We investigate some geometric properties of the real algebraic variety Δ Δ Δ of symmetric matrices with repeated eigenvalues. We explicitly compute the volume of its intersection with the sphere and prove a Eckart-Young-Mirsky-type theorem for the distance function from a generic matrix to points in Δ Δ Δ . We exhibit conne…
Study of repeated games with unobserved agent rewards using MAB framework.
problem Designing policies for principals in repeated principal-agent games with unobservable agent rewards.
method Developed a policy achieving low regret (square-root regret up to a log factor) for perfect-knowledge agents.
result Constructed an estimator for agent's expected reward and designed a policy achieving low regret.
New algorithm reduces online learning regret for bounded recall games.
problem Reducing regret in online learning with limited past information.
method Constructing a stationary bounded-recall algorithm with O ( 1 / M ) O(1/\sqrt{M}) O ( 1/ M ) regret. result Any low regret bounded-recall algorithm must be aware of past losses' order.
Study shows how adaptive market agents can lead to persistent overpricing in financial markets.
problem Persistent overpricing in financial markets by adaptive market agents.
method Analyzes a repeated game between market maker and market taker, decomposes the game into competitive and collaborative components, and uses projected stochastic gradient ascent.
result Decentralized learning by adaptive market agents can lead to persistent overpricing in financial markets.
The paper analyzes how mutable blockchain protocols affect miner behavior and strategic stability.
problem The mutability of blockchain protocols undermines long-term planning and cooperative equilibria.
method Integrates Austrian capital theory with repeated game theory to examine miner behavior under different institutional conditions.
result Effective time preference increases when protocol rules are mutable, leading to political rent-seeking and undermining strategic coherence.
Human behavioural patterns exhibit selfish or competitive, as well as selfless or altruistic tendencies, both of which have demonstrable effects on human social and economic activity. In behavioural economics, such effects have traditionally been illustrated experimentally via simple games like the dictator and ultimat…
GAME improves matrix completion by considering subgroup-specific latent structures.
problem Heterogeneous data with overlapping categories, smoothing away subgroup-specific variation.
method Group-Aware Matrix Estimation (GAME) with overlapping nuclear-norm penalties.
result GAME outperforms global low-rank estimators in structured missingness regimes.
Study on HFTs' interactions with a large trader using mean field game theory.
problem Interactions between high-frequency traders and a large trader executing assets at discrete times.
method Modeling HFTs' behavior using a jump process and solving the equilibrium through mean field game approach.
result Inventory-averse HFTs lower LT's costs when market impact is large.
New algorithm reduces risk in online games with limited feedback.
problem Risk-averse learning in repeated unknown games with bandit feedback.
method Proposes a momentum-based algorithm to estimate CVaR using historical cost values.
result Achieves sub-linear regret and outperforms existing methods in numerical experiments.
New approach for adaptive conformal inference using Blackwell's theory.
problem Non-exchangeable environments in sequential conformal inference.
method Reinterpretation of ACI as a game, construction of coverage and efficiency objectives, approachability strategy.
result Algorithm achieves strong theoretical guarantees and practical insights.
New method accelerates smooth games using spectral shape analysis.
problem Accelerating optimization in smooth games with complex numerical challenges.
method Matrix iteration theory and spectral shape analysis to characterize and manipulate acceleration.
result Identified a continuum of optimization strategies from convex minimization to gradient descent.
Paper develops a classification method using matrix-variate t-distributions.
problem Classifying matrix-valued observations with dependence structure.
method Develops an Expectation-Maximization algorithm for discriminant analysis.
result Method shows promise on various datasets.
Study shows market makers can cooperate without communication.
problem Concerns of collusion in AI-driven market-making.
method Formulated as a repeated game, studied with Q-learning.
result Market makers can learn cooperative strategies without communication.
New algorithms minimize regret with global costs in online learning.
problem Minimizing regret in online learning with global costs.
method Extended FTRL algorithms for Blackwell's approachability.
result First bounds on regret minimization with explicit dependence in p p p and d d d . NeuPL learns diverse policies in strategy games efficiently.
problem Iterative training of policies in strategy games leads to under-trained good-responses and wasteful repetition.
method NeuPL uses a single conditional model to represent a population of policies, offering convergence guarantees and transfer learning.
result NeuPL achieves better performance and efficiency across various domains, enabling access to novel strategies.
The paper introduces a frequency-domain estimator for low-order systems from noisy data.
problem Estimating frequency responses of low-order systems from noisy measurements.
method Uses a quadratic data-fitting term regularized by the nuclear norm of a Loewner matrix, subject to a convex stability constraint.
result Proves a finite-sample error bound and extends it to all frequencies through rational interpolation.
New algorithms achieve logarithmic regret in KL-regularized Markov games.
problem Improving sample efficiency in game-theoretic settings with KL regularization.
method Developed OMG and SOMG algorithms for matrix and Markov games, using best response sampling and superoptimistic bonuses.
result Logarithmic regret in T T T that scales inversely with KL regularization strength β β β . Study of repeated principal-agent bandit game with self-interested and exploratory learning agents.
problem Interaction between principal and agent in unknown environments with learning and exploration behaviors.
method Developed algorithms for self-interested and exploratory learning agents with bandit feedback, achieving regret bounds.
result Achieved O ~ ( T 2 / 3 ) \widetilde{O}(T^{2/3}) O ( T 2/3 ) regret bound for exploratory learning agent in i.i.d. reward setup. New measure of policy regret shows compatibility with traditional external regret in adversarial games.
problem Incompatibility between traditional and new policy regret measures in adaptive adversaries.
method Revisited policy regret and compared it with external regret; introduced policy equilibrium.
result Policy regret and external regret are compatible in adversarial games.
User strategization undermines algorithmic trustworthiness.
problem User strategic behavior corrupts algorithmic data and trust.
method Modeling user-platform interactions as a game, analyzing strategic behavior's short-term benefits and long-term harms.
result User strategization can initially benefit platforms but ultimately harms their ability to make accurate decisions.
Extends trading framework to incorporate real-world constraints.
problem Trading strategies in multi-player non-cooperative games with constraints.
method Re-framed as quadratic programming problem, constraints readily incorporated.
result Two-trader equilibria calculated dynamically.
Generative adversarial networks (GANs) are successful deep generative models. GANs are based on a two-player minimax game. However, the objective function derived in the original motivation is changed to obtain stronger gradients when learning the generator. We propose a novel algorithm that repeats the density ratio e…
New algorithm tackles adversarial bandits with budget constraints.
problem Adversarial multi-armed bandits with supply/budget constraints.
method Design of an algorithm with O(log T) competitive ratio.
result Achieved a competitive ratio of O(log T) for the adversarial version.
New framework values football players based on in-game interactions.
problem Valuing football players based on in-game performance.
method Combining financial models and network theory using a passing matrix.
result Dynamic and individualized player valuation framework.
We introduce a novel Bayesian hybrid matrix factorisation model (HMF) for data integration, based on combining multiple matrix factorisation methods, that can be used for in- and out-of-matrix prediction of missing values. The model is very general and can be used to integrate many datasets across different entity type…
We consider the problem of influence maximization in fixed networks for contagion models in an adversarial setting. The goal is to select an optimal set of nodes to seed the influence process, such that the number of influenced nodes at the conclusion of the campaign is as large as possible. We formulate the problem as…