The paper formalizes and analyzes multi-agent Q-learning with value factorization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper studies optimal approximation factors in misspecified off-policy RL, identifying key factors under various settings.
Optimizes risk measures given known marginal distributions of two unknown factors.
This paper introduces a new framework for learning nearly decomposable Q-functions via communication minimization.
Long term optimal investment problems are studied in a factor model with matrix valued state variables. Explicit parameter restrictions are obtained under which, for an isoelastic investor, the finite horizon value function and optimal strategy converge to their long-run counterparts as the investment horizon approache…
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones f…
In distributed function computation, each node has an initial value and the goal is to compute a function of these values in a distributed manner. In this paper, we propose a novel token-based approach to compute a wide class of target functions to which we refer as "Token-based function Computation with Memory" (TCM) …
N-discount optimality was introduced as a hierarchical form of policy- and value-function optimality, with Blackwell optimality lying at the top level of the hierarchy Veinott (1969); Blackwell (1962). We formalize notions of myopic discount factors, value functions and policies in terms of Blackwell optimality in MDPs…
Study market-to-book ratios using Stochastic Portfolio Theory.
Study optimal investment and consumption in a stochastic factor model.
We propose a model for the credit markets in which the random default times of bonds are assumed to be given as functions of one or more independent "market factors". Market participants are assumed to have partial information about each of the market factors, represented by the values of a set of market factor informa…
In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…
Authors improve accuracy analysis for portfolio optimization with multiple timescale factors.
We consider Chern-Simons theory with complex gauge group and present a complete non-perturbative evaluation of the path integral (the partition function and certain expectation values of Wilson loops) on Seifert fibred 3-Manifolds. We use the method of Abelianisation. In certain cases the path integral can be seen to f…
By Gyongy's theorem, a local and stochastic volatility (LSV) model is calibrated to the market prices of all European call options with positive maturities and strikes if its local volatility function is equal to the ratio of the Dupire local volatility function over the root conditional mean square of the stochastic v…
Matrix factorization is a well-studied task in machine learning for compactly representing large, noisy data. In our approach, instead of using the traditional concept of matrix rank, we define a new notion of link-rank based on a non-linear link function used within factorization. In particular, by applying the round …
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
In this paper we assume the insurance wealth process is driven by the compound Poisson process. The discounting factor is modelled as a geometric Brownian motion at first and then as an exponential function of an integrated Ornstein-Uhlenbeck process. The objective is to maximize the cumulated value of expected discoun…
A new RL paradigm reduces state-action-value function approximation inefficiency.
Proposes a nonparametric tensor factorization for sparse data.
We consider derivative-free algorithms for stochastic and non-stochastic convex optimization problems that use only function values rather than gradients. Focusing on non-asymptotic bounds on convergence rates, we show that if pairs of function values are available, algorithms for -dimensional optimization that use …
Novel KAN-based autoencoder improves asset pricing models' accuracy and interpretability.
New bound on Rademacher complexity for vector functions.
In an incomplete market, with incompleteness stemming from stochastic factors imperfectly correlated with the underlying stocks, we derive representations of homothetic (power, exponential and logarithmic) forward performance processes in factor-form using ergodic BSDE. We also develop a connection between the forward …
Paper proposes an analytical pricing model for puttable bonds with credit risk.
Policy gradient methods with aggregated states can achieve better performance than approximate policy iteration.
Pricing formulae for defaultable corporate bonds with discrete coupons under consideration of the government taxes in the united model of structural and reduced form models are provided. The aim of this paper is to generalize the comprehensive structural model for defaultable fixed income bonds (considered in [1]) into…
GPLVMF improves CARS performance by addressing overfitting and context importance.
In this paper we investigate the curvature of conformal deformations by noncommutative Weyl factors of a flat metric on a noncommutative 2-torus, by analyzing in the framework of spectral triples functionals associated to perturbed Dolbeault operators. The analogue of Gaussian curvature turns out to be a sum of two fun…
The article calculates a multiplying factor to convert rational Vassiliev invariants to integer-valued ones.
Enhances FM models for numerical features using function basis encoding.
We consider an incomplete market with a nontradable stochastic factor and a continuous time investment problem with an optimality criterion based on monotone mean-variance preferences. We formulate it as a stochastic differential game problem and use Hamilton-Jacobi-Bellman-Isaacs equations to find an optimal investmen…
Temporal difference learning explained through gradient splitting, improving convergence times.
We study the optimal excess-of-loss reinsurance problem when both the intensity of the claims arrival process and the claim size distribution are influenced by an exogenous stochastic factor. We assume that the insurer's surplus is governed by a marked point process with dual-predictable projection affected by an envir…
Method regularizes Cholesky factors to detect nonstationarity in longitudinal data.
Proposes CC-NMDF for analyzing manifold-valued data.
The Shapley value theory is used for risk allocation in non-orthogonal risk factors.
We factorize harmonic maps with values in a semisimple Lie groups in a product of harmonic maps with values in the components of the Iwasawa decomposition. In particular, we use this factorization to study the harmonic maps from into .
We extend the calculus of adiabatic pseudo-differential operators to study the adiabatic limit behavior of the eta and zeta functions of a differential operator , constructed from an elliptic family of operators indexed by . We show that the regularized values and are smooth functions of …
A new method STMF improves missing value prediction using tropical semiring.
Holomorphic functions from knot complements link to quantum modular forms.
We explain the main concepts of Prospect Theory and Cumulative Prospect Theory within the framework of rational dynamic asset pricing theory. We derive option pricing formulas when asset returns are altered with a generalized Prospect Theory value function or a modified Prelec weighting probability function and introdu…
Stochastic discount factor (SDF) processes in dynamic economies admit a permanent-transitory decomposition in which the permanent component characterizes pricing over long investment horizons. This paper introduces an empirical framework to analyze the permanent-transitory decomposition of SDF processes. Specifically, …
This paper presents a sequential randomized lowrank matrix factorization approach for incrementally predicting values of an unknown function at test points using the Gaussian Processes framework. It is well-known that in the Gaussian processes framework, the computational bottlenecks are the inversion of the (regulariz…
New algorithm speeds up NMF with -divergence.
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run. Yet, it may be difficult (or even intractable) mathematically to learn with th…
Proposes a VAE variant for ordinal content factors.
In this article, we consider a 2 factors-model for pricing defaultable bond with discrete default intensity and barrier where the 2 factors are stochastic risk free short rate process and firm value process. We assume that the default event occurs in an expected manner when the firm value reaches a given default barrie…