Policy gradient methods find Nash equilibrium in noisy games.
problem Finding Nash equilibrium in noisy games.
method Policy gradient methods with noise added.
result Policy gradient methods converge to Nash equilibrium in noisy games.
Develops a reinforcement learning algorithm for learning deterministic equilibrium policies in time-inconsistent control problems.
problem Learning equilibrium policies in time-inconsistent control problems.
method Continuous-time model-free reinforcement learning algorithm using deterministic policy gradient approach.
result Learned equilibrium policies in general time-inconsistent control problems.
A new approach to MV portfolio optimization with jumps and RL.
problem Continuous-time Mean-Variance portfolio optimization with jumps.
method Jump-diffusion process, Reinforcement Learning, time-inconsistent control (TIC), Actor-Critic RL algorithm.
result The proposed RL model is profitable in real-world market data.
Algorithm learns Nash equilibria in stochastic games using entropy-regularized policies.
problem Learning Nash equilibria in zero-sum stochastic games is computationally expensive.
method Entropy-regularized soft policies for Q-function updates.
result Algorithm converges to Nash equilibrium under certain conditions.
Equilibrium pricing has been proven to underlie the rational Insured expectancy of premia additivity for composition of policies fully covering independent risks.
Naive investors make riskier choices than optimal strategies in continuous-time finance.
problem Continuous-time Markowitz portfolio selection with naive reoptimization.
method Analytical derivation of naive policies from discretely naive policies.
result Naive policies are always riskier and less efficient than equilibrium policies.
We consider the problem of finding Nash equilibrium for two-player turn-based zero-sum games. Inspired by the AlphaGo Zero (AGZ) algorithm, we develop a Reinforcement Learning based approach. Specifically, we propose Explore-Improve-Supervise (EIS) method that combines "exploration", "policy improvement"' and "supervis…
Policy mirror ascent achieves Nash equilibrium in mean field games without a population generative model.
problem Achieving Nash equilibrium in mean field games without a population generative model.
method Policy mirror ascent, contractive operator, single-path TD learning.
result Policy mirror ascent converges to Nash equilibrium within O ~ ( ε − 2 ) \widetilde{\mathcal{O}}(\varepsilon^{-2}) O ( ε − 2 ) samples. Paper improves AIRL by enhancing policy imitation and addressing reward recovery issues.
problem Inadequate policy imitation and limited transferable reward recovery in AIRL.
method Substituted built-in algorithm with SAC for policy updating and proposed PPO-AIRL + SAC hybrid framework.
result SAC improves policy imitation but hinders reward recovery; PPO-AIRL + SAC achieves satisfactory transfer effect.
Entropy regularization improves MFG learning efficiency and stability.
problem Improving Mean Field Game learning efficiency and stability.
method Entropy regularization applied to MFG with learning.
result Entropy regularization yields time-dependent policies and stabilizes convergence.
Dynamic pricing policy converges to Nash equilibrium with low regret.
problem Sequential price competition among sellers over multiple periods.
method Semi-parametric least-squares estimation of s-concave demand functions.
result Prices converge to Nash equilibrium with rate O ( T − 1 / 7 ) O(T^{-1/7}) O ( T − 1/7 ) and sellers incur regret O ( T 5 / 7 ) O(T^{5/7}) O ( T 5/7 ) . Pessimistic Minimax Value Iteration finds efficient NE policies from offline data.
problem Finding an approximate Nash equilibrium in offline Markov games with non-uniform coverage.
method Pessimistic Minimax Value Iteration (PMVI) constructs pessimistic value function estimates and solves NEs.
result Established a nearly minimax optimal result for offline Markov games with function approximation.
Study optimal treatment assignment policies under strategic agent responses.
problem Learning optimal treatment policies with strategic agents complicates estimation.
method Dynamic model with threshold convergence to mean-field equilibrium, consistent estimator for policy gradient.
result Threshold for treatment assignment converges to mean-field equilibrium threshold under large but finite number of agents.
Conventional economic analysis of stringent climate change mitigation policy generally concludes various levels of economic slowdown as a result of substantial spending on low carbon technology. Equilibrium economics however could not explain or predict the current economic crisis, which is of financial nature. Meanwhi…
Study efficient offline RL in Markov games with general models.
problem Learn approximate equilibria from offline data in Markov games.
method Use Bellman-consistent pessimism for interval estimation and optimize gap relaxation.
result First framework for sample-efficient offline learning in Markov games, handling all equilibria.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
Policy gradient method proves convergence in imperfect-information games.
problem Policy gradient methods in imperfect-information games (EFGs).
method Policy gradient approach with best-iterate convergence.
result Policy gradient leads to provable best-iterate convergence in self-play EFGs.
Study proposes new OPE estimators for two-player zero-sum games.
problem Evaluating new policies using historical data from a different policy in multi-player zero-sum games.
method Doubly robust and double reinforcement learning estimators to project exploitability.
result Prove exploitability estimation error bounds and regret bounds for policy profiles.
Study examines how governance, corruption, and R&D affect economic development.
problem The impact of corruption and governance on economic development.
method General equilibrium model with heterogeneous agents and a government, including corruption as a fraction of tax revenues.
result Redistribution and innovation-led strategies can mitigate the negative effects of corruption on economic development.
DQN outperforms static policies in a dynamic fee environment for automated market makers.
problem How automated market makers (AMMs) perform under dynamic fees is unknown.
method Constructed a closed-loop simulator with dynamic fees, noise flow, and arbitrage.
result A small DQN policy outperforms static policies in a dynamic fee environment.
The paper analyzes a game where players must balance short-term and long-term interests, leading to cooperative or competitive outcomes.
problem Analyzing time inconsistency in inter-personal decision-making under non-exponential discounting.
method Iterative procedures and Zorn's lemma to find Nash equilibria between players' intra-personal equilibria.
result Inter-personal equilibria exist and depend on the impatience levels of the players.
This paper considers the Merton portfolio management problem. We are concerned with non-exponential discounting of time and this leads to time inconsistencies of the decision maker. Following Ekeland and Pirvu 2006, we introduce the notion of equilibrium policies and we characterize them by an integral equation. The ma…
In a continuous time stochastic economy, this paper considers the problem of consumption and investment in a financial market in which the representative investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switches…
UBI model proves financial equilibrium exists.
problem Proving existence of financial equilibrium with UBI.
method Backward stochastic differential equation (BSDE) approach.
result Equilibrium exists in UBI model.
Model shows how financial markets can decarbonize under climate uncertainty.
problem Decarbonization of financial markets under climate uncertainty.
method Mean-field game approach to model firm decisions and investor interactions.
result Climate uncertainty weakens the impact of green-minded investors on decarbonization.
The notion of \emph{policy regret} in online learning is a well defined? performance measure for the common scenario of adaptive adversaries, which more traditional quantities such as external regret do not take into account. We revisit the notion of policy regret and first show that there are online learning settings …
New RL algorithms learn QSE from strategic feedbacks with sample efficiency.
problem Learning QSE in Markov games with strategic feedbacks.
method Proposes sample-efficient algorithms for online and offline settings, combining quantal response model learning and RL.
result Achieves sublinear regret bounds and quantifies model uncertainty.
A new approach to fine-tuning LLMs with human feedback.
problem Inability of current reward models to fully represent human preferences.
method Introducing NLHF, a new pipeline for LLM fine-tuning using pairwise human feedback.
result NLHF produces a sequence of policies converging to the regularized Nash equilibrium.
An unconventional approach for optimal stopping under model ambiguity is introduced. Besides ambiguity itself, we take into account how ambiguity-averse an agent is. This inclusion of ambiguity attitude, via an α α α -maxmin nonlinear expectation, renders the stopping problem time-inconsistent. We look for subgame perfect…
This paper addresses the problem of multi-agent inverse reinforcement learning (MIRL) in a two-player general-sum stochastic game framework. Five variants of MIRL are considered: uCS-MIRL, advE-MIRL, cooE-MIRL, uCE-MIRL, and uNE-MIRL, each distinguished by its solution concept. Problem uCS-MIRL is a cooperative game in…
New algorithm improves self-play reinforcement learning for competitive games.
problem Inefficient opponent selection in self-play reinforcement learning.
method Intelligently selects opponents based on adversarial rules derived from saddle point optimization.
result Algorithm converges to approximate equilibrium with high probability in convex-concave games.
Study compares employers with and without anticipating strategic labor force responses.
problem Understanding and optimizing strategic interactions in labor markets.
method Formulation of causal strategic classification, theory, and experiments.
result Performatively optimal hiring policies improve employer and labor outcomes, but can also harm labor force utility.
Paper presents a GMFG framework for large stochastic games.
problem Learning Nash Equilibrium in large stochastic games.
method Value-based and policy-based reinforcement learning algorithms with smoothed policies.
result Proposed algorithms GMF-V and GMF-P are efficient and robust in GMFG setting.
This paper outlines a critical gap in the assessment methodology used to estimate the macroeconomic costs and benefits of climate policy. It shows that the vast majority of models used for assessing climate policy use assumptions about the financial system that sit at odds with the observed reality. In particular, the …
A new approach for cooperative multi-agent reinforcement learning with limited communication, reducing the number of communication rounds.
problem Limited communication in decentralized MARL systems leads to outdated information and unstable learning.
method Base policy prediction technique to estimate gradients and collect samples for a sequence of base policies.
result The proposed algorithm converges to an ε-Nash equilibrium with significantly fewer communication rounds and samples.
Investor optimizes portfolio under dynamic risk preferences.
problem Optimizing investment under uncertain future risk attitudes.
method Developed a general equilibrium framework and solved for subgame-perfect equilibrium policies.
result Equilibrium policies include a novel hedging component to counteract anticipated risk aversion changes.
Study dynamic asset allocation in incomplete markets using game theory and nonlocal BSDEs.
problem Dynamic mean-variance asset allocation in general incomplete markets with non-exponential discounting.
method Game-theoretic approach, decomposition into myopic and hedging strategies, nonlocal BSDEs, fixed-point theorem.
result Well-posedness of solutions to BSDEs, existence of equilibrium control policy.
This paper simplifies complex game dynamics by using a recursive representation.
problem Difficulties in finite-player dynamic games with private information.
method Provides a recursive representation and noise-state model.
result Equilibrium becomes a deterministic fixed point in impulse-response functions.
New RL algorithms find SNE in Markov games with myopic followers.
problem Finding SNE in Markov games with myopic followers.
method Optimistic and pessimistic variants of least-squares value iteration, incorporating function approximation.
result First provably efficient RL algorithms for SNEs in general-sum Markov games with myopic followers.
A new method for RLHF using proximal point Nash learning.
problem Capturing real human preferences in RLHF.
method Proximal point Nash learning, embedding self-play updates into a proximal point framework.
result High-probability last-iterate convergence for the combined method.
A method to learn robust policies for environments with model mismatches.
problem Training agents in high-stakes scenarios with mismatched training and real environments.
method Formalizes the perturbation as a zero-sum game to find Nash Equilibrium, which corresponds to the robust policy.
result Our algorithm can find a near-optimal robust policy with high probability using polynomial samples.
DORIS algorithm achieves no-regret learning in Markov games with adversarial opponents.
problem Decentralized policy learning in Markov games with nonstationary opponents.
method DORIS algorithm using optimistic hyperpolicy mirror descent.
result Achieves K \sqrt{K} K -regret in general function approximation. Models analyze strategic risk-taking in continuous action games.
problem Strategic risk-taking dynamics in continuous action games.
method Normal form game, multi-player scenarios, regret minimization algorithms, numerical algorithm for calculation.
result Nash equilibrium also serves as a correlated equilibrium in continuous games.
New algorithm finds Nash equilibrium in multi-agent games.
problem Extending RL theory to multi-agent settings with large state spaces.
method Proposes an algorithm using an exploiter to find Nash equilibrium policies.
result Proves finding Nash equilibrium is possible with polynomial samples.
Algorithm finds ε-equilibrium policies for multi-agent Markov games with hidden low-rank structure.
problem Designing efficient algorithms for multi-agent Markov games with unknown representation and hidden low-rank structure.
method Model-based and model-free approaches using representation learning to construct an effective representation from data.
result Achieves poly ( H , d , A , 1 / ε ) (H,d,A,1/\varepsilon) ( H , d , A , 1/ ε ) sample complexity for both model-based and model-free approaches. This paper tackles learning Stackelberg equilibrium in asymmetric games efficiently from noisy samples.
problem Learning Stackelberg equilibrium in asymmetric, general-sum games efficiently from noisy samples.
method The paper initiates the theoretical study of sample-efficient learning of the Stackelberg equilibrium in bandit feedback setting.
result Sharp positive results on sample-efficient learning of Stackelberg equilibrium with value optimal up to a fundamental gap identified.
General equilibrium is the dominant theoretical framework for economic policy analysis at the level of the whole economy. In practice, general equilibrium treats economies as being always in equilibrium, albeit in a sequence of equilibria as driven by external changes in parameters. This view is sometimes defended on t…
Paper introduces metrics for evaluating multi-agent policies using best response dynamics.
problem Evaluation and ranking of multi-agent policies in reinforcement learning.
method Adopting strict best response dynamics (SBRD) to model selfish behaviors, proposing perturbed SBRD for dynamic and non-stationary settings.
result Proposed perturbed SBRD can observe policies with maximum metrics and differ from optimal by any given tolerance.