Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

23466891 · Jun 202019922001200920172026
48 results for Quantal response equilibrium

New RL algorithms learn QSE from strategic feedbacks with sample efficiency.

problem Learning QSE in Markov games with strategic feedbacks.
method Proposes sample-efficient algorithms for online and offline settings, combining quantal response model learning and RL.
result Achieves sublinear regret bounds and quantifies model uncertainty.

Experimental economics has repeatedly demonstrated that the Nash equilibrium makes inaccurate predictions for a vast set of games. Instead, several alternative theoretical concepts predict behavior that is much more in tune with observed data, with the quantal response equilibrium as the most prominent example. However…

2010-12-03abs ↗pdf ↗

New framework recovers reward and rationality parameters from game behavior.

problem Statistical ambiguity in identifying reward and rationality parameters in competitive games.
method Blind Inverse Game Theory (Blind-IGT) using entropy-regularized Quantal Response Equilibrium and Normalized Least Squares (NLS) estimator.
result Optimal convergence rate of O(N1/2)\mathcal{O}(N^{-1/2}) for joint parameter recovery.

Unified framework for estimating reward functions in competitive games.

problem Estimating unknown reward functions in competitive games.
method Unified framework with entropy regularization for reward function recovery.
result Strong theoretical guarantees and practical effectiveness demonstrated.

Study compares cryptocurrency and stock markets using statistical equilibrium models.

problem Comparing the stochastic structure of cryptocurrency and stock markets.
method Applied QRSE model to analyze daily returns of cryptocurrencies and S&P 500 companies.
result Revealed differences in informational efficiency between cryptocurrency and stock markets.

Model captures decision-making under bounded rationality with prior beliefs and market feedback.

problem Bounded rationality in decision-making with limited processing abilities.
method Maximum entropy principle applied to Quantal Response Statistical Equilibrium framework.
result Prior beliefs influence decision-making, altering the outcome of market feedback.

The dual crises of the sub-prime mortgage crisis and the global financial crisis has prompted a call for explanations of non-equilibrium market dynamics. Recently a promising approach has been the use of agent based models (ABMs) to simulate aggregate market dynamics. A key aspect of these models is the endogenous emer…

2018-09-05abs ↗pdf ↗

Paper generalizes Hardy-Rogers maps for market equilibrium analysis in duopoly markets.

problem Existence and uniqueness of market equilibrium in duopoly markets with non-differentiable, nonlinear response functions.
method Coupled fixed points approach for generalized Hardy-Rogers maps.
result Enriched understanding of market equilibrium in duopoly markets with non-differentiable response functions.

This paper extends a Kyle model to include price-responsive traders, revealing new dynamics and equilibria.

problem Real-world market dynamics involve price-responsive traders, affecting market equilibrium and insider profits.
method Developed a continuous-time Kyle model with two types of price-responsive traders (momentum and contrarian), leading to a forward-backward Riccati system for equilibrium.
result The model shows that feedback effects can lead to multiple equilibria and amplify price informativeness.

Paper studies zero-sum games with noisy observations and identifies equilibrium conditions.

problem Zero-sum games with noisy observations of the leader's actions.
method Analyzes the equilibrium of games with noisy action observability, identifies necessary conditions for uniqueness, and investigates the cardinality of best responses.
result The noisy observations significantly impact the cardinality of the follower's set of best responses, and under certain conditions, this set becomes a singleton almost surely.

New CGMD model predicts non-equilibrium processes better than existing methods.

problem Inconsistency in conditional distribution of unresolved variables.
method Time-lagged independent component analysis to minimize entropy contribution of unresolved variables.
result The model's generalization ability for non-equilibrium processes is significantly improved.

In this paper, we propose an equilibrium pricing model in a dynamic multi-period stochastic framework with uncertain income streams. In an incomplete market, there exist two traded risky assets (e.g. stock/commodity and weather derivative) and a non-traded underlying (e.g. temperature). The risk preferences are of expo…

2012-05-28abs ↗pdf ↗

New algorithm for solving minimax problems over distributions converges to Nash equilibrium.

problem Solving minimax problems over probability distributions.
method Symmetric Mean-field Langevin Dynamics (MFL-AG and MFL-ABR) with weighted averaging and best response dynamics.
result Converges to mixed Nash equilibrium with average-iterate and last-iterate convergence.

Study compares employers with and without anticipating strategic labor force responses.

problem Understanding and optimizing strategic interactions in labor markets.
method Formulation of causal strategic classification, theory, and experiments.
result Performatively optimal hiring policies improve employer and labor outcomes, but can also harm labor force utility.

Study best-response learning dynamics in zero-sum polymatrix games under full and minimal information settings.

problem Learning dynamics in zero-sum polymatrix games under different information settings.
method Two-timescale learning dynamics combining smoothed best-response updates and TD-learning for estimating local payoff functions.
result Polynomial-time finite-sample guarantees for convergence to an ε-Nash equilibrium in the minimal information case.

The large majority of risk-sharing transactions involve few agents, each of whom can heavily influence the structure and the prices of securities. This paper proposes a game where agents' strategic sets consist of all possible sharing securities and pricing kernels that are consistent with Arrow-Debreu sharing rules. F…

2014-12-13abs ↗pdf ↗

Paper introduces metrics for evaluating multi-agent policies using best response dynamics.

problem Evaluation and ranking of multi-agent policies in reinforcement learning.
method Adopting strict best response dynamics (SBRD) to model selfish behaviors, proposing perturbed SBRD for dynamic and non-stationary settings.
result Proposed perturbed SBRD can observe policies with maximum metrics and differ from optimal by any given tolerance.

Investors' strategic trading affects asset prices, modeled as a game.

problem Investors' trading rates influence asset prices in dynamic markets.
method Model as a non-zero sum singular stochastic differential game, establishing equivalence between best-response and auxiliary control problems.
result Unique Nash equilibrium is deterministic with a closed-form solution.

Study examines how market dynamics affect emissions trading prices and abatement efforts.

problem Effectiveness of emissions markets depends on regulatory standards, costs, and abatement levels.
method Radner equilibrium framework that considers intertemporal decision-making and uncertainty.
result Variations in regulatory standards, costs, and abatement levels influence allowance prices and abatement efforts.

The paper analyzes a game where players must balance short-term and long-term interests, leading to cooperative or competitive outcomes.

problem Analyzing time inconsistency in inter-personal decision-making under non-exponential discounting.
method Iterative procedures and Zorn's lemma to find Nash equilibria between players' intra-personal equilibria.
result Inter-personal equilibria exist and depend on the impatience levels of the players.

New RL algorithms find SNE in Markov games with myopic followers.

problem Finding SNE in Markov games with myopic followers.
method Optimistic and pessimistic variants of least-squares value iteration, incorporating function approximation.
result First provably efficient RL algorithms for SNEs in general-sum Markov games with myopic followers.

Model analyzes trading frictions in cap-and-trade markets, showing how they interact to affect market effectiveness.

problem Analyzing how trading frictions impact cap-and-trade market effectiveness.
method Developed a dynamic stochastic model with multiple trading frictions, characterized access choices in closed form, and quantified using EU ETS data.
result Trading frictions interact to amplify or dampen market responses, and their combined effect is non-additive.

Informed traders strategically reveal noisier signals, making prices less responsive to public information.

problem How informed traders strategically reveal signals impacts market prices and utility.
method Modeling a market with an informed trader, an uninformed trader, and liquidity providers, proving equilibrium existence.
result In equilibrium, the insider strategically reveals a noisier signal, making prices less responsive to public information.

We address the challenge of designing optimal adversarial noise algorithms for settings where a learner has access to multiple classifiers. We demonstrate how this problem can be framed as finding strategies at equilibrium in a two-player, zero-sum game between a learner and an adversary. In doing so, we illustrate the…

2019-06-06abs ↗pdf ↗

Study efficient offline RL in Markov games with general models.

problem Learn approximate equilibria from offline data in Markov games.
method Use Bellman-consistent pessimism for interval estimation and optimize gap relaxation.
result First framework for sample-efficient offline learning in Markov games, handling all equilibria.

Develops variational framework for LQG risk-sensitive MFGs with major-minor interactions.

problem Risk-sensitive optimal control in LQG systems with major-minor interactions.
method Variational approach, nonlinear necessary and sufficient condition of optimality, equivalent risk-neutral measure, Markovian closed-loop best-response strategies.
result Derives optimal control strategies for LQG risk-sensitive MFGs with major-minor interactions, establishing Nash and ε\varepsilon-Nash equilibria.

Policy mirror ascent achieves Nash equilibrium in mean field games without a population generative model.

problem Achieving Nash equilibrium in mean field games without a population generative model.
method Policy mirror ascent, contractive operator, single-path TD learning.
result Policy mirror ascent converges to Nash equilibrium within O~(ε2)\widetilde{\mathcal{O}}(\varepsilon^{-2}) samples.

Paper explores limits and possibilities of aligning LLMs with human preferences.

problem Aligning LLMs with diverse human preferences to ensure fairness and informed outcomes.
method Analysis of probabilistic representation of human preferences and preservation of diverse preferences.
result LLMs can't fully align with human preferences using reward-based approaches due to Condorcet cycles, but mixed strategies are statistically possible.

During the last two years, Europe has been facing a debt crisis, and Greece has been at its center. In response to the crisis, drastic actions have been taken, including the halving of Greek debt. Policy makers acted because interest rates for sovereign debt increased dramatically. High interest rates imply that defaul…

2012-09-27abs ↗pdf ↗

Study of portfolio management under relative performance concerns using mean field games.

problem Portfolio management problems under relative performance concerns.
method Forward utilities of CARA type, mean field games, best response and equilibrium strategies.
result Solve forward-utility finite player game and mean-field game under asset specialization.

We present the quantum model of Bertrand duopoly and study the entanglement behavior on the profit functions of the firms. Using the concept of optimal response of each firm to the price of the opponent, we found only one Nash equilibirum point for maximally entangled initial state. The very presence of quantum entangl…

2010-01-16abs ↗pdf ↗

Study shows climate change can cause a 'run on fossil fuels' affecting prices and production.

problem Impact of climate change expectations on fossil fuel markets and prices.
method Dynamic, general equilibrium model of climate-change-linked transition risk.
result Climate change expectations can lead to either increased or decreased fossil fuel prices, depending on economic responses.

SPPO optimizes language model alignment by treating preferences as a game and achieving state-of-the-art performance.

problem Capturing intransitivity and irrationality in human preferences for accurate language model alignment.
method Self-play-based approach to identify Nash equilibrium policy through iterative policy updates.
result SPPO achieves state-of-the-art win-rate of 28.53% on AlpacaEval 2.0 without external supervision.

The paper analyzes how investors' wealth can decline collectively under partial information.

problem Investors' wealth can decline collectively under partial information.
method The paper derives a Nash equilibrium for mean-variance portfolio selection under relative performance criteria, considering both full and partial information.
result Relative performance criteria can lead to downward self-reinforcement of investors' wealth, which is more pronounced under partial information.

A multi-layer deep Gaussian process (DGP) model is a hierarchical composition of GP models with a greater expressive power. Exact DGP inference is intractable, which has motivated the recent development of deterministic and stochastic approximation methods. Unfortunately, the deterministic approximation methods yield a…

2019-10-26abs ↗pdf ↗

Analyzing real data on international trade covering the time interval 1950-2000, we show that in each year over the analyzed period the network is a typical representative of the ensemble of maximally random weighted networks, whose directed connections (bilateral trade volumes) are only characterized by the product of…

2011-04-13abs ↗pdf ↗

Study time-inconsistent portfolio optimization for competitive agents with relative performance criteria.

problem Time-inconsistent mean field and n-agent games under relative performance criteria.
method Construct open-loop equilibrium strategies for n-agent games and mean field games.
result Explicit solutions for n-agent games and mean field games, unique in a special class of equilibria.

Paper studies optimal tracking portfolio in mean field game of large fund competition.

problem Optimal tracking portfolio in large fund competition with relative performance benchmark.
method Formulated mean field game problem, established existence of mean field equilibrium using PDE approach, constructed approximate Nash equilibrium.
result Existence of mean field equilibrium and consistency condition verified.