Introduces expected eligibility traces for more efficient credit assignment in reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Policy evaluation with linear function approximation is an important problem in reinforcement learning. When facing high-dimensional feature spaces, such a problem becomes extremely hard considering the computation efficiency and quality of approximations. We propose a new algorithm, LSTD()-RP, which leverages rando…
New methods improve deep reinforcement learning by accelerating credit assignment.
Meta-learning adjusts TD learning's eligibility trace parameter for more efficient reinforcement learning.
Recently, a new multi-step temporal learning algorithm, called , unifies -step Tree-Backup (when ) and -step Sarsa (when ) by introducing a sampling parameter . However, similar to other multi-step temporal-difference learning algorithms, needs much memory consumption and computation tim…
New algorithms for collaborative reinforcement learning with limited communication.
Off-policy reinforcement learning with eligibility traces is challenging because of the discrepancy between target policy and behavior policy. One common approach is to measure the difference between two policies in a probabilistic way, such as importance sampling and tree-backup. However, existing off-policy learning …
Reinforcement learning (RL) has had many successes in both "deep" and "shallow" settings. In both cases, significant hyperparameter tuning is often required to achieve good performance. Furthermore, when nonlinear function approximation is used, non-stationarity in the state representation can lead to learning instabil…
This paper motivates and develops source traces for temporal difference (TD) learning in the tabular setting. Source traces are like eligibility traces, but model potential histories rather than immediate ones. This allows TD errors to be propagated to potential causal states and leads to faster generalization. Source …
Full-sampling (e.g., Q-learning) and pure-expectation (e.g., Expected Sarsa) algorithms are efficient and frequently used techniques in reinforcement learning. Q is the first approach unifies them with eligibility trace through the sampling degree . However, it is limited to the tabular case, for large-scale …
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement learning, its theoretical analysis has proved challenging and few guarantees on its s…
Temporal-difference (TD) learning is an important field in reinforcement learning. Sarsa and Q-Learning are among the most used TD algorithms. The Q() algorithm (Sutton and Barto (2017)) unifies both. This paper extends the Q() algorithm to an online multi-step algorithm Q() using eligibility traces and int…
Within the context of capital adequacy, we study comonotonicity of risk measures in terms of the primitives of the theory: acceptance sets and eligible, or reference, assets. We show that comonotonicity cannot be characterized by the properties of the acceptance set alone and heavily depends on the choice of the eligib…
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…
Estimates treatment effects in bipartite systems with partial eligibility and interference.
We study capital requirements for bounded financial positions defined as the minimum amount of capital to invest in a chosen eligible asset targeting a pre-specified acceptability test. We allow for general acceptance sets and general eligible assets, including defaultable bonds. Since the payoff of these assets is not…
We discuss risk measures representing the minimum amount of capital a financial institution needs to raise and invest in a pre-specified eligible asset to ensure it is adequately capitalized. Most of the literature has focused on cash-additive risk measures, for which the eligible asset is a risk-free bond, on the grou…
The risk of financial positions is measured by the minimum amount of capital to raise and invest in eligible portfolios of traded assets in order to meet a prescribed acceptability constraint. We investigate nondegeneracy, finiteness and continuity properties of these risk measures with respect to multiple eligible ass…
This study improves convergence of two-timescale SA under Markovian noise in reinforcement learning.
Interventional cancer clinical trials are generally too restrictive, and some patients are often excluded on the basis of comorbidity, past or concomitant treatments, or the fact that they are over a certain age. The efficacy and safety of new treatments for patients with these characteristics are, therefore, not defin…
In real-world applications of reinforcement learning (RL), noise from inherent stochasticity of environments is inevitable. However, current policy evaluation algorithms, which plays a key role in many RL algorithms, are either prone to noise or inefficient. To solve this issue, we introduce a novel policy evaluation a…
Effective complements to human judgment, artificial intelligence techniques have started to aid human decisions in complicated social problems across the world. In the context of United States for instance, automated ML/DL classification models offer complements to human decisions in determining Medicaid eligibility. H…
We show that the linear trace Harnack quadratic on a steady gradient Ricci soliton satisfies the heat equation. Similar result holds for shrinkers. We also present an interpolation between Perelman's and Cao--Hamilton's Harnacks on a steady soliton.
Bayesian method corrects timing misalignment in recurrent event studies.
Modern deep reinforcement learning methods have departed from the incremental learning required for eligibility traces, rendering the implementation of the -return difficult in this context. In particular, off-policy methods that utilize experience replay remain problematic because their random sampling of minibatch…
A dynamic Boltzmann machine (DyBM) has been proposed as a model of a spiking neural network, and its learning rule of maximizing the log-likelihood of given time-series has been shown to exhibit key properties of spike-timing dependent plasticity (STDP), which had been postulated and experimentally confirmed in the fie…
Ray tracing sampler improves neural network sampling efficiency and resilience.
This work develops a fully decentralized multi-agent algorithm for policy evaluation. The proposed scheme can be applied to two distinct scenarios. In the first scenario, a collection of agents have distinct datasets gathered following different behavior policies (none of which is required to explore the full state spa…
In a capital adequacy framework, risk measures are used to determine the minimal amount of capital that a financial institution has to raise and invest in a portfolio of pre-specified eligible assets in order to pass a given capital adequacy test. From a capital efficiency perspective, it is important to identify the s…
A new method for optimizing deep neural networks using TKFAC.
Differentiable relaxation for inferring partial orders from noisy linear data.
Novel approach finds implicit regularisation in two-player games using BEA.
Monetary risk measures are usually interpreted as the smallest amount of external capital that must be added to a financial position to make it acceptable. We propose a new concept: intrinsic risk measures and argue that this approach provides a direct path from unacceptable positions towards the acceptance set. Intrin…
Study gradient flow of phase transitions with fixed contact angle.
One of the crucial problems in mathematical finance is to mitigate the risk of a financial position by setting up hedging positions of eligible financial securities. This leads to focusing on set-valued maps associating to any financial position the set of those eligible payoffs that reduce the risk of the position to …
We introduce a method called TracIn that computes the influence of a training example on a prediction made by the model. The idea is to trace how the loss on the test point changes during the training process whenever the training example of interest was utilized. We provide a scalable implementation of TracIn via: (a)…
Early training phase affects deep neural network optimization and generalization.
Conditions for trivial gradient hyperbolic Ricci and Yamabe solitons to be Einstein or constant scalar curvature.
Algorithmic fairness involves expressing notions such as equity, or reasonable treatment, as quantifiable measures that a machine learning algorithm can optimise. Most work in the literature to date has focused on classification problems where the prediction is categorical, such as accepting or rejecting a loan applica…
NGD models have higher effective dimension than SGD models.
Study proves structure results for homogeneous spaces supporting specific equations.
Paper estimates differences in multi-attribute Gaussian graphical models using non-convex penalties.
Improved DP-SGD for variational inference reduces noise and variance.
Set-valued risk measures on with for conical market models are defined, primal and dual representation results are given. The collection of initial endowments which allow to super-hedge a multivariate claim are shown to form the values of a set-valued sublinear (coherent) risk measure. Sc…
TRACE analyzes risk changes in models trained on shifted data.
This paper investigates robust and efficient DR/RDR estimators for WATEs.
In this paper, we discuss the isometric embedding problem in hyperbolic space with nonnegative extrinsic curvature. We prove a priori bounds for the trace of the second fundamental form H and extend the result to n-dimensions. We also obtain an estimate for the gradient of the smaller principal curvature in 2 dimension…
Zeroth-order methods favor flat minima in machine learning.