Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

12243648 · May 202619922001200920172026
48 results for eligibility traces

Introduces expected eligibility traces for more efficient credit assignment in reinforcement learning.

problem Efficiently assigning credit to states and actions in reinforcement learning.
method Introduces expected eligibility traces, allowing updates to counterfactual sequences.
result Substantial improvements in temporal-difference learning can be achieved with expected traces.

Meta-learning adjusts TD learning's eligibility trace parameter for more efficient reinforcement learning.

problem Efficiently tuning the eligibility trace parameter for temporal difference learning.
method Meta-learning method to adjust eligibility trace parameter state-dependently.
result Improves overall quality of update targets, minimizing target error.

This paper motivates and develops source traces for temporal difference (TD) learning in the tabular setting. Source traces are like eligibility traces, but model potential histories rather than immediate ones. This allows TD errors to be propagated to potential causal states and leads to faster generalization. Source …

2019-02-08abs ↗pdf ↗

New methods improve deep reinforcement learning by accelerating credit assignment.

problem Challenges in achieving fast and stable off-policy learning in deep reinforcement learning.
method Extends the generalized PBE objective to support multistep credit assignment and derives three gradient-based methods.
result Proposed methods outperform PPO and StreamQ in MuJoCo and MinAtar environments.

New algorithms for collaborative reinforcement learning with limited communication.

problem Efficiently learning value functions in multi-agent systems with strict information constraints.
method Distributed gradient-based temporal difference algorithms with consensus schemes.
result Parameter estimates converge to ODEs with defined invariant sets under general assumptions.

We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…

2016-02-16abs ↗pdf ↗

Estimates treatment effects in bipartite systems with partial eligibility and interference.

problem Randomized experiments in bipartite systems with partial treatment eligibility and interference.
method Formalizes eligibility-constrained bipartite experiments, defines PTTE and STTE, identifies conditions, develops ensemble estimators, introduces projection.
result Proposed estimators recover PTTE and STTE with low bias and variance, corrects interference bias in field experiments.

We study capital requirements for bounded financial positions defined as the minimum amount of capital to invest in a chosen eligible asset targeting a pre-specified acceptability test. We allow for general acceptance sets and general eligible assets, including defaultable bonds. Since the payoff of these assets is not…

2012-03-20abs ↗pdf ↗

We discuss risk measures representing the minimum amount of capital a financial institution needs to raise and invest in a pre-specified eligible asset to ensure it is adequately capitalized. Most of the literature has focused on cash-additive risk measures, for which the eligible asset is a risk-free bond, on the grou…

2012-06-03abs ↗pdf ↗

The risk of financial positions is measured by the minimum amount of capital to raise and invest in eligible portfolios of traded assets in order to meet a prescribed acceptability constraint. We investigate nondegeneracy, finiteness and continuity properties of these risk measures with respect to multiple eligible ass…

2013-08-15abs ↗pdf ↗

Monetary risk measures are usually interpreted as the smallest amount of external capital that must be added to a financial position to make it acceptable. We propose a new concept: intrinsic risk measures and argue that this approach provides a direct path from unacceptable positions towards the acceptance set. Intrin…

2016-10-27abs ↗pdf ↗

This study improves convergence of two-timescale SA under Markovian noise in reinforcement learning.

problem Stability and convergence of two-timescale stochastic approximations under Markovian noise.
method Introduced a new control strategy for the fast timescale parameter.
result Established almost sure convergence of TDC with eligibility traces under off-policy learning with linear function approximation.

Algorithmic fairness involves expressing notions such as equity, or reasonable treatment, as quantifiable measures that a machine learning algorithm can optimise. Most work in the literature to date has focused on classification problems where the prediction is categorical, such as accepting or rejecting a loan applica…

2020-01-16abs ↗pdf ↗

Modern deep reinforcement learning methods have departed from the incremental learning required for eligibility traces, rendering the implementation of the λλ-return difficult in this context. In particular, off-policy methods that utilize experience replay remain problematic because their random sampling of minibatch…

2018-10-23abs ↗pdf ↗

Set-valued risk measures on LdpL^p_d with 0p0 \leq p \leq \infty for conical market models are defined, primal and dual representation results are given. The collection of initial endowments which allow to super-hedge a multivariate claim are shown to form the values of a set-valued sublinear (coherent) risk measure. Sc…

2010-11-27abs ↗pdf ↗

This paper investigates robust and efficient DR/RDR estimators for WATEs.

problem Lack of systematic investigation into robustness and efficiency conditions for WATE estimation.
method Proposes three RDR estimators using semiparametric efficient influence function and double/debiased machine learning.
result Demonstrates the practical relevance of the methods in medical and social sciences.

Develops a risk score to assist ECMO planning for critically ill patients with viral or unspecified pneumonia.

problem Lack of a risk score to guide ECMO planning for critically ill patients.
method Leverages machine learning to develop the PEER score.
result PEER score predicts mortality and decompensation in patients eligible for ECMO.

Derives Selberg trace formula on Riemann surfaces and generalizes to other spaces.

problem Deriving and generalizing the Selberg trace formula.
method Supersymmetric localization principle and path integral derivation.
result Derives Selberg trace formula on arbitrary compact Riemann surfaces and generic compact locally symmetric spaces.

New complete panel dataset for LMICs helps analyze innovation and development.

problem Lack of complete data for empirical analyses in LMICs.
method Predictive Mean Matching multiple imputation technique.
result Created a large dataset of 47 variables for 82 LMICs from 2005-2019.

CausalSim corrects bias in trace-driven simulations for more accurate results.

problem Bias in trace-driven simulations due to system conditions during trace collection.
method CausalSim learns a causal model of system dynamics and latent factors from an RCT to remove bias from trace data.
result CausalSim reduces simulation errors by 53% and 61% compared to baselines, providing more accurate insights.

This paper investigates the strength of the trace field as a commensurability invariant of hyperbolic 3-manifolds. We construct an infinite family of two-component hyperbolic link complements which are pairwise incommensurable and have the same trace field, and infinitely many 1-cusped finite volume hyperbolic 3-manifo…

2007-08-08abs ↗pdf ↗