Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

231462693924 · Jun 202019922001200920172026
48 results for value function factorization

The paper formalizes and analyzes multi-agent Q-learning with value factorization.

problem Understanding and improving the convergence of multi-agent Q-learning with value factorization.
method Formalized a multi-agent fitted Q-iteration framework for analyzing factorized multi-agent Q-learning.
result Multi-agent Q-learning with linear value factorization can converge under certain conditions.

This paper studies optimal approximation factors in misspecified off-policy RL, identifying key factors under various settings.

problem Understanding optimal approximation factors in misspecified off-policy value function estimation.
method Examined various settings including weighted L2L_2-norm, LL_\infty norm, state aliasing, and state coverage.
result Established optimal asymptotic approximation factors for different norms and identified two instance-dependent factors for L2(μ)L_2(μ) norm.

Optimizes risk measures given known marginal distributions of two unknown factors.

problem Determining an upper bound for spectral risk measures with unknown joint distribution.
method Introduces Maximum Spectral Measure (MSP) as a worst-case risk measure, formulated as an optimization problem with a more general objective function.
result Characterizes the continuity properties of the optimal value function and optimal solution set with respect to marginal distributions.

This paper introduces a new framework for learning nearly decomposable Q-functions via communication minimization.

problem Challenges in multi-agent reinforcement learning, especially scalability and non-stationarity.
method Learning nearly decomposable Q-functions (NDQ) via communication minimization, introducing two information-theoretic regularizers.
result Significantly outperforms baseline methods on the StarCraft unit micromanagement benchmark.

Long term optimal investment problems are studied in a factor model with matrix valued state variables. Explicit parameter restrictions are obtained under which, for an isoelastic investor, the finite horizon value function and optimal strategy converge to their long-run counterparts as the investment horizon approache…

2014-08-29abs ↗pdf ↗

In distributed function computation, each node has an initial value and the goal is to compute a function of these values in a distributed manner. In this paper, we propose a novel token-based approach to compute a wide class of target functions to which we refer as "Token-based function Computation with Memory" (TCM) …

2017-03-26abs ↗pdf ↗

Study optimal investment and consumption in a stochastic factor model.

problem Optimal investment and consumption decisions in a stochastic factor model.
method Characterization of well-posedness, numerical algorithm, and general theory of sub- and supersolutions for HJB equation.
result Proves existence and provides bounds for the solution to the HJB equation.

We propose a model for the credit markets in which the random default times of bonds are assumed to be given as functions of one or more independent "market factors". Market participants are assumed to have partial information about each of the market factors, represented by the values of a set of market factor informa…

2010-06-15abs ↗pdf ↗

In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…

2019-06-24abs ↗pdf ↗

Authors improve accuracy analysis for portfolio optimization with multiple timescale factors.

problem Asymptotic accuracy of portfolio optimization approximations for general utility functions and two timescale factors.
method Construct sub- and super-solutions to fully nonlinear problem.
result Rigorous justification of accuracy for portfolio optimization with general utility functions and two timescale factors.

Matrix factorization is a well-studied task in machine learning for compactly representing large, noisy data. In our approach, instead of using the traditional concept of matrix rank, we define a new notion of link-rank based on a non-linear link function used within factorization. In particular, by applying the round …

2018-05-01abs ↗pdf ↗

Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…

2019-02-19abs ↗pdf ↗

A new RL paradigm reduces state-action-value function approximation inefficiency.

problem Challenges in state-action-value function approximation for RL.
method State Action Separable Reinforcement Learning (sasRL) decouples action space from value function learning.
result sasRL achieves up to 75% better performance than state-of-the-art MDP-based RL algorithms.

Proposes a nonparametric tensor factorization for sparse data.

problem Handling sparse tensor data with structural and interpretability benefits.
method Hierarchical Gamma processes and Poisson random measures for tensor-valued process, Dirichlet processes for sampling entry indices, Gaussian processes for values.
result Demonstrates superior performance on benchmark datasets.

Paper proposes an analytical pricing model for puttable bonds with credit risk.

problem Analytical pricing of puttable bonds with credit risk.
method Developed a 2-factor structural PDE model and derived analytical pricing formula under specific conditions.
result Derived analytical pricing formula for puttable bonds with credit risk.

Policy gradient methods with aggregated states can achieve better performance than approximate policy iteration.

problem Approximation errors in policy and value function approximations.
method State-aggregated representations and policy gradient methods.
result Policy gradient methods can achieve a per-period regret bounded by ε, while approximate policy iteration and value iteration have a higher regret.

Pricing formulae for defaultable corporate bonds with discrete coupons under consideration of the government taxes in the united model of structural and reduced form models are provided. The aim of this paper is to generalize the comprehensive structural model for defaultable fixed income bonds (considered in [1]) into…

2013-09-06abs ↗pdf ↗

GPLVMF improves CARS performance by addressing overfitting and context importance.

problem Overfitting and lack of automatic context importance determination in GP-based CARS.
method GPLVMF applies a non-zero mean function and real-valued latent space to improve GP model performance.
result Significant improvement in performance on real datasets and automatic context importance determination.

In this paper we investigate the curvature of conformal deformations by noncommutative Weyl factors of a flat metric on a noncommutative 2-torus, by analyzing in the framework of spectral triples functionals associated to perturbed Dolbeault operators. The analogue of Gaussian curvature turns out to be a sum of two fun…

2011-10-16abs ↗pdf ↗

The article calculates a multiplying factor to convert rational Vassiliev invariants to integer-valued ones.

problem Converting rational valued Vassiliev invariants to integer-valued ones.
method Calculates the minimal multiplying factor λ needed for rational Vassiliev invariants to become integer-valued.
result Obtains a set of integer-valued Vassiliev invariants.

Enhances FM models for numerical features using function basis encoding.

problem Challenges in incorporating numerical features into FM variants.
method Encoding numerical features into a vector of function values for learning segmentized functions.
result Improves model accuracy by learning segmentized functions of numerical features.

Temporal difference learning explained through gradient splitting, improving convergence times.

problem Learning value functions in Markov Decision Processes with linear approximations.
method Interpreting TD learning as gradient splitting and applying convergence proofs from gradient descent.
result Improved convergence times for TD learning, especially with a minor variation.

We study the optimal excess-of-loss reinsurance problem when both the intensity of the claims arrival process and the claim size distribution are influenced by an exogenous stochastic factor. We assume that the insurer's surplus is governed by a marked point process with dual-predictable projection affected by an envir…

2019-04-10abs ↗pdf ↗

Method regularizes Cholesky factors to detect nonstationarity in longitudinal data.

problem Detecting nonstationarity in large covariance matrices of longitudinal data.
method Fused-Lasso regularization on Cholesky factors.
result Regularization leads to smooth subdiagonals, indicating nonstationarity.

The Shapley value theory is used for risk allocation in non-orthogonal risk factors.

problem Risk allocation among non-orthogonal risk factors in financial portfolios.
method Using Shapley value from cooperative game theory to allocate risk contributions.
result Explicit formulas and numerical algorithms for calculating risk allocations are derived.

We factorize harmonic maps with values in a semisimple Lie groups in a product of harmonic maps with values in the components of the Iwasawa decomposition. In particular, we use this factorization to study the harmonic maps from Rn\mathbb{R}^n into SL(2,R)SL(2,\mathbb{R}).

2015-06-15abs ↗pdf ↗

We extend the calculus of adiabatic pseudo-differential operators to study the adiabatic limit behavior of the eta and zeta functions of a differential operator δδ, constructed from an elliptic family of operators indexed by S1S^1. We show that the regularized values η(δt,0)η(δ_t,0) and tζ(δt,0)tζ(δ_t,0) are smooth functions of …

2002-04-12abs ↗pdf ↗

A new method STMF improves missing value prediction using tropical semiring.

problem Limited capability of linear models to model complex relations.
method Sparse Tropical Matrix Factorization (STMF) using tropical semiring.
result STMF outperforms NMF on real data, especially in handling extreme values.

Stochastic discount factor (SDF) processes in dynamic economies admit a permanent-transitory decomposition in which the permanent component characterizes pricing over long investment horizons. This paper introduces an empirical framework to analyze the permanent-transitory decomposition of SDF processes. Specifically, …

2014-12-15abs ↗pdf ↗

In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run. Yet, it may be difficult (or even intractable) mathematically to learn with th…

2019-02-05abs ↗pdf ↗

Proposes a VAE variant for ordinal content factors.

problem Isolating ordinal-valued content factors in deep latent variable models.
method Introduces a partially ordered set (poset) structure and a conditional Gaussian spacing prior model.
result Significant improvements in content-style separation over previous non-ordinal approaches.