Study on mean field games with singular controls and their applications.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on investment strategy for agents with periodic preferences and discounting.
In this paper, we settle the sampling complexity of solving discounted two-player turn-based zero-sum stochastic games up to polylogarithmic factors. Given a stochastic game with discount factor we provide an algorithm that computes an -optimal strategy with high-probability given $\tilde{O}((1 - γ)^{-3}…
We consider a discounted reward control problem in continuous time stochastic environment where the discount rate might be an unbounded function of the control process. We provide a set of general assumptions to ensure that there exists a smooth classical solution to the corresponding HJB equation. Moreover, some verif…
The paper analyzes a game where players must balance short-term and long-term interests, leading to cooperative or competitive outcomes.
The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…
Study time-inconsistent portfolio optimization for competitive agents with relative performance criteria.
Study optimal portfolio strategies with time-varying discount rates.
We describe an approximate dynamic programming (ADP) approach to compute approximations of the optimal strategies and of the minimal losses that can be guaranteed in discounted repeated games with vector-valued losses. Such games prominently arise in the analysis of regret in repeated decision-making in adversarial env…
Paper analyzes robust strategies in a pension plan game with ambiguous financial markets.
Investment decisions shift earlier as patience decreases, with implications for pasting conditions.
We develop an option pricing model based on a tug-of-war game. This two-player zero-sum stochastic differential game is formulated in the context of a multi-dimensional financial market. The issuer and the holder try to manipulate asset price processes in order to minimize and maximize the expected discounted reward. W…
In this paper we revisit the method of off-policy corrections for reinforcement learning (COP-TD) pioneered by Hallak et al. (2017). Under this method, online updates to the value function are reweighted to avoid divergence issues typical of off-policy learning. While Hallak et al.'s solution is appealing, it cannot ea…
We propose a simple model of the banking system incorporating a game feature where the evolution of monetary reserve is modeled as a system of coupled Feller diffusions. The Markov Nash equilibrium generated through minimizing the linear quadratic cost subject to Cox-Ingersoll-Ross type processes creates liquidity and …
We propose an analytically tractable variation of the minority game in which rational agents use probabilistic strategies. In our model, agents choose between two alternatives repeatedly, and those who are in the minority get a pay-off 1, others zero. The agents optimize the expectation value of their discounted fu…
Multiplayer Online Battle Arena (MOBA) is currently one of the most popular genres of digital games around the world. The domain of knowledge contained in these complicated games is large. It is hard for humans and algorithms to evaluate the real-time game situation or predict the game result. In this paper, we introdu…
The paper analyzes Q-learning in 2-player Markov games and provides gap-dependent logarithmic regret bounds.
We determine the optimal strategy for investing in a Black-Scholes market in order to maximize the probability that wealth at death meets a bequest goal , a type of goal-seeking problem, as pioneered by Dubins and Savage (1965, 1976). The individual consumes at a constant rate , so the level of wealth required fo…
We study the problem of super-replication for game options under proportional transaction costs. We consider a multidimensional continuous time model, in which the discounted stock price process satisfies the conditional full support property. We show that the super-replication price is the cheapest cost of a trivial s…
Algorithm converges to Nash equilibria in competitive games.
Consider a two-player zero-sum stochastic game where the transition function can be embedded in a given feature space. We propose a two-player Q-learning algorithm for approximating the Nash equilibrium strategy via sampling. The algorithm is shown to find an -optimal strategy using sample size linear to the number …
Study competitive energy markets using stochastic impulse games.
Pessimistic model-based algorithm finds Nash equilibria in zero-sum Markov games from offline data.
We consider the problem of finding stationary Nash equilibria (NE) in a finite discounted general-sum stochastic game. We first generalize a non-linear optimization problem from Filar and Vrieze [2004] to a -player setting and break down this problem into simpler sub-problems that ensure there is no Bellman error fo…
Investor and firm optimize sustainable investment and emission reduction through a dynamic game.
Despite significant advances in the field of deep Reinforcement Learning (RL), today's algorithms still fail to learn human-level policies consistently over a set of diverse tasks such as Atari 2600 games. We identify three key challenges that any algorithm needs to master in order to perform well on all games: process…
This paper optimizes model-based RL for two-player zero-sum games with near-optimal sample complexity.
Paper proposes a mean-field gradient descent for zero-sum games, proving convergence to Nash equilibrium.
A new method for reinforcement learning scales errors without tuning.
Study on reinsurance decisions using mean-variance criterion with irreversible contracts.
In reinforcement learning, Return, which is the weighted accumulated future rewards, and Value, which is the expected return, serve as the objective that guides the learning of the policy. In classic RL, return is defined as the exponentially discounted sum of future rewards. One key insight is that there could be many…
We examine two different techniques for parameter averaging in GAN training. Moving Average (MA) computes the time-average of parameters, whereas Exponential Moving Average (EMA) computes an exponentially discounted sum. Whilst MA is known to lead to convergence in bilinear settings, we provide the -- to our knowledge …
Study optimal investment-reinsurance strategies in equity-linked insurance products using Stackelberg game theory.
Optimization of parameterized policies for reinforcement learning (RL) is an important and challenging problem in artificial intelligence. Among the most common approaches are algorithms based on gradient ascent of a score function representing discounted return. In this paper, we examine the role of these policy gradi…
Study solves HJB equations for time-inconsistent control problems.
Model shows how financial markets can decarbonize under climate uncertainty.
New approach to optimal dividend control with mean-variance criterion.
Paper develops a discounted algorithm for online convex optimization that adapts to unknown discount factors.
Paper introduces non-linear discounting models for default compensation and climate valuation.
Study dynamic asset allocation in incomplete markets using game theory and nonlocal BSDEs.
Study analyzes how discounts affect train ticket purchases and rescheduling in Switzerland.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
Proves lower discount rates are needed for future losses.
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
There is an observed basis between repo discounting, implied from market repo rates, and bond discounting, stripped from the market prices of the underlying bonds. Here, this basis is explained as a convexity effect arising from the decorrelation between the discount rates for derivatives and bonds. Using a Hull-White …
Proposes a new framework for discount models.
When we implement a portfolio selection methodology under a mean-risk formulation, it is essential to correctly model investors' risk aversion which may be time-dependent, or even state-dependent during the investment procedure. In this paper, we propose a behavior risk aversion model, which is a piecewise linear funct…
New method handles large reward variations in reinforcement learning.