Introduces expected eligibility traces for more efficient credit assignment in reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Policy evaluation with linear function approximation is an important problem in reinforcement learning. When facing high-dimensional feature spaces, such a problem becomes extremely hard considering the computation efficiency and quality of approximations. We propose a new algorithm, LSTD()-RP, which leverages rando…
Meta-learning adjusts TD learning's eligibility trace parameter for more efficient reinforcement learning.
Recently, a new multi-step temporal learning algorithm, called , unifies -step Tree-Backup (when ) and -step Sarsa (when ) by introducing a sampling parameter . However, similar to other multi-step temporal-difference learning algorithms, needs much memory consumption and computation tim…
Off-policy reinforcement learning with eligibility traces is challenging because of the discrepancy between target policy and behavior policy. One common approach is to measure the difference between two policies in a probabilistic way, such as importance sampling and tree-backup. However, existing off-policy learning …
This paper motivates and develops source traces for temporal difference (TD) learning in the tabular setting. Source traces are like eligibility traces, but model potential histories rather than immediate ones. This allows TD errors to be propagated to potential causal states and leads to faster generalization. Source …
New methods improve deep reinforcement learning by accelerating credit assignment.
New algorithms for collaborative reinforcement learning with limited communication.
Temporal-difference (TD) learning is an important field in reinforcement learning. Sarsa and Q-Learning are among the most used TD algorithms. The Q() algorithm (Sutton and Barto (2017)) unifies both. This paper extends the Q() algorithm to an online multi-step algorithm Q() using eligibility traces and int…
Within the context of capital adequacy, we study comonotonicity of risk measures in terms of the primitives of the theory: acceptance sets and eligible, or reference, assets. We show that comonotonicity cannot be characterized by the properties of the acceptance set alone and heavily depends on the choice of the eligib…
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…
Estimates treatment effects in bipartite systems with partial eligibility and interference.
Reinforcement learning (RL) has had many successes in both "deep" and "shallow" settings. In both cases, significant hyperparameter tuning is often required to achieve good performance. Furthermore, when nonlinear function approximation is used, non-stationarity in the state representation can lead to learning instabil…
We study capital requirements for bounded financial positions defined as the minimum amount of capital to invest in a chosen eligible asset targeting a pre-specified acceptability test. We allow for general acceptance sets and general eligible assets, including defaultable bonds. Since the payoff of these assets is not…
We discuss risk measures representing the minimum amount of capital a financial institution needs to raise and invest in a pre-specified eligible asset to ensure it is adequately capitalized. Most of the literature has focused on cash-additive risk measures, for which the eligible asset is a risk-free bond, on the grou…
The risk of financial positions is measured by the minimum amount of capital to raise and invest in eligible portfolios of traded assets in order to meet a prescribed acceptability constraint. We investigate nondegeneracy, finiteness and continuity properties of these risk measures with respect to multiple eligible ass…
Full-sampling (e.g., Q-learning) and pure-expectation (e.g., Expected Sarsa) algorithms are efficient and frequently used techniques in reinforcement learning. Q is the first approach unifies them with eligibility trace through the sampling degree . However, it is limited to the tabular case, for large-scale …
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement learning, its theoretical analysis has proved challenging and few guarantees on its s…
Interventional cancer clinical trials are generally too restrictive, and some patients are often excluded on the basis of comorbidity, past or concomitant treatments, or the fact that they are over a certain age. The efficacy and safety of new treatments for patients with these characteristics are, therefore, not defin…
In real-world applications of reinforcement learning (RL), noise from inherent stochasticity of environments is inevitable. However, current policy evaluation algorithms, which plays a key role in many RL algorithms, are either prone to noise or inefficient. To solve this issue, we introduce a novel policy evaluation a…
Effective complements to human judgment, artificial intelligence techniques have started to aid human decisions in complicated social problems across the world. In the context of United States for instance, automated ML/DL classification models offer complements to human decisions in determining Medicaid eligibility. H…
Bayesian method corrects timing misalignment in recurrent event studies.
A dynamic Boltzmann machine (DyBM) has been proposed as a model of a spiking neural network, and its learning rule of maximizing the log-likelihood of given time-series has been shown to exhibit key properties of spike-timing dependent plasticity (STDP), which had been postulated and experimentally confirmed in the fie…
This work develops a fully decentralized multi-agent algorithm for policy evaluation. The proposed scheme can be applied to two distinct scenarios. In the first scenario, a collection of agents have distinct datasets gathered following different behavior policies (none of which is required to explore the full state spa…
In a capital adequacy framework, risk measures are used to determine the minimal amount of capital that a financial institution has to raise and invest in a portfolio of pre-specified eligible assets in order to pass a given capital adequacy test. From a capital efficiency perspective, it is important to identify the s…
Monetary risk measures are usually interpreted as the smallest amount of external capital that must be added to a financial position to make it acceptable. We propose a new concept: intrinsic risk measures and argue that this approach provides a direct path from unacceptable positions towards the acceptance set. Intrin…
This study improves convergence of two-timescale SA under Markovian noise in reinforcement learning.
One of the crucial problems in mathematical finance is to mitigate the risk of a financial position by setting up hedging positions of eligible financial securities. This leads to focusing on set-valued maps associating to any financial position the set of those eligible payoffs that reduce the risk of the position to …
Algorithmic fairness involves expressing notions such as equity, or reasonable treatment, as quantifiable measures that a machine learning algorithm can optimise. Most work in the literature to date has focused on classification problems where the prediction is categorical, such as accepting or rejecting a loan applica…
Modern deep reinforcement learning methods have departed from the incremental learning required for eligibility traces, rendering the implementation of the -return difficult in this context. In particular, off-policy methods that utilize experience replay remain problematic because their random sampling of minibatch…
Set-valued risk measures on with for conical market models are defined, primal and dual representation results are given. The collection of initial endowments which allow to super-hedge a multivariate claim are shown to form the values of a set-valued sublinear (coherent) risk measure. Sc…
This paper investigates robust and efficient DR/RDR estimators for WATEs.
This is a continuation of our previous work arXiv:1601.05617 on trace and inverse trace of Steklov eigenvalues. More new inequalities for the trace and inverse trace of Steklov eigenvalues are obtained.
Study of torus surgeries on knot traces, finding exotic surfaces and traces.
Classifies knot traces with specific trisection genus limits.
Guillemin trace formula adapted for group actions.
Paper derives trace formula for magnetic Laplacian at zero energy.
In this paper, we obtain some new estimates for the trace and inverse trace of Steklov eigenvalues. The estimates generalize some previous results of Hersch-Payne-Schiffer , Brock}, Raulot-Savo and Dittmar.
Develops a risk score to assist ECMO planning for critically ill patients with viral or unspecified pneumonia.
New methods derive a generalized Frenkel trace formula for Lie groups.
Examines a new type of analytic torsion on Riemannian manifolds.
Derives Selberg trace formula on Riemann surfaces and generalizes to other spaces.
Introduces a new model for mapping matrices to matrices, subsuming linear regression.
New complete panel dataset for LMICs helps analyze innovation and development.
CausalSim corrects bias in trace-driven simulations for more accurate results.
New findings on knots and their traces, distinguishing L-space knots by their 0-trace.
This paper investigates the strength of the trace field as a commensurability invariant of hyperbolic 3-manifolds. We construct an infinite family of two-component hyperbolic link complements which are pairwise incommensurable and have the same trace field, and infinitely many 1-cusped finite volume hyperbolic 3-manifo…
We give axioms which characterize the local Reidemeister trace for orientable differentiable manifolds. The local Reidemeister trace in fixed point theory is already known, and we provide both uniqueness and existence results for the local Reidemeister trace in coincidence theory.