The true online TD(λ) algorithm has recently been proposed (van Seijen and Sutton, 2014) as a universal replacement for the popular TD(λ) algorithm, in temporal-difference learning and reinforcement learning. True online TD(λ) has better theoretical properties than conventional TD(λ), and the expectation is that it als…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
RO-TD learns sparse value functions efficiently.
Introduces TDRC to balance TD's ease and soundness.
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side. In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms and provide theoretical analysis on their conver…
Improved TD learning reduces batch sampling error.
Uniform TD(0) bound derived for function approximation with Markov noise.
New TD algorithms stabilize RL tasks by reformulating updates into fixed point equations.
New TD method stabilizes average-reward learning.
Improved TD(0) algorithm for reinforcement learning with linear approximations.
In reinforcement learning, the TD() algorithm is a fundamental policy evaluation method with an efficient online implementation that is suitable for large-scale problems. One practical drawback of TD() is its sensitivity to the choice of the step-size. It is an empirically well-known fact that a large step-size l…
Improved TD learning reduces variance and bias errors.
Our understanding of reinforcement learning (RL) has been shaped by theoretical and empirical results that were obtained decades ago using tabular representations and linear function approximators. These results suggest that RL methods that use temporal differencing (TD) are superior to direct Monte Carlo estimation (M…
Improved TD learning with tail averaging and regularization achieves optimal convergence rates.
In this paper, we introduce a method for adapting the step-sizes of temporal difference (TD) learning. The performance of TD methods often depends on well chosen step-sizes, yet few algorithms have been developed for setting the step-size automatically for TD learning. An important limitation of current methods is that…
Paper proposes an active multi-step TD algorithm for reinforcement learning.
The family of temporal difference (TD) methods span a spectrum from computationally frugal linear methods like TD(λ) to data efficient least squares methods. Least square methods make the best use of available data directly computing the TD solution and thus do not require tuning a typically highly sensitive learning r…
The study proposes using TD error for selecting σ in Q(σ, λ).
Study on distributional TD learning with linear approximations for better return estimation.
While there are convergence guarantees for temporal difference (TD) learning when using linear function approximators, the situation for nonlinear models is far less understood, and divergent examples are known. Here we take a first step towards extending theoretical convergence guarantees to TD learning with nonlinear…
We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (TD) learning and Monte Carlo (MC) policy evaluation. The two methods are known to achieve complementary bias-variance trade-off properties, w…
Unified framework for finite-sample RL algorithms using Lyapunov theory.
The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy improvement scheme. However, most Temporal Difference (TD) based solutions ignore the discrepancy between the stationary distribution of the beh…
In this paper, we provide a unified analysis of temporal difference learning algorithms with linear function approximators by exploiting their connections to Markov jump linear systems (MJLS). We tailor the MJLS theory developed in the control community to characterize the exact behaviors of the first and second order …
TD learning reduces interference, leading to better generalization.
Accelerates TD learning for long-horizon reinforcement learning problems.
Temporal-difference (TD) networks are a class of predictive state representations that use well-established TD methods to learn models of partially observable dynamical systems. Previous research with TD networks has dealt only with dynamical systems with finite sets of observations and actions. We present an algorithm…
Paper shows TD learning without projection converges robustly.
Maps from buildings to spaces study K-theory of Hecke algebras.
Paper improves TD(0) convergence rate with LFA, i.i.d. samples, and averaging.
TDS provides exact samples for conditional distributions in diffusion models.
Proposes a convergent TD algorithm for off-policy RL.
Stochastic differential equation approximation for linear TD(0) under Markovian noise
Paper analyzes TD() convergence rates for arbitrary features.
TDprop uses Jacobi preconditioning to improve adaptive optimizers in Deep RL.
Paper presents a technique using Spearman's Rank Correlation Coefficient for KE in TDs.
New algorithms improve distributional TD learning with linear approximations.
In unsupervised domain adaptation (UDA), classifiers for the target domain (TD) are trained with clean labeled data from the source domain (SD) and unlabeled data from TD. However, in the wild, it is difficult to acquire a large amount of perfectly clean labeled data in SD given limited budget. Hence, we consider a new…
TD-Flow improves long-term predictions in agent learning.
TD(0) with Polyak-Ruppert averaging achieves robust and fast convergence rates
Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coupling leads to nonconvexity and even divergence in optimization. As a result, the global convergence …
We present a database of parliamentary debates that contains the complete record of parliamentary speeches from Dáil Éireann, the lower house and principal chamber of the Irish parliament, from 1919 to 2013. In addition, the database contains background information on all TDs (Teachta Dála, members of parliament), such…
We provide non-asymptotic bounds for the well-known temporal difference learning algorithm TD(0) with linear function approximators. These include high-probability bounds as well as bounds in expectation. Our analysis suggests that a step-size inversely proportional to the number of iterations cannot guarantee optimal …
Temporal Difference Learning analysis under non-i.i.d. data and nonlinear approximation.
Q()-Learning improves Q-Learning by separating action-value functions into different time scales.
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process. Although TD is one of the most widely used algorithms in reinforcement learning, its theoretical analysis has proved challenging and few guarantees on its s…
In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …
Temporal-difference (TD) learning is an important field in reinforcement learning. Sarsa and Q-Learning are among the most used TD algorithms. The Q() algorithm (Sutton and Barto (2017)) unifies both. This paper extends the Q() algorithm to an online multi-step algorithm Q() using eligibility traces and int…
Improved TD learning with neural nets reduces sample complexity and overparameterization.