A new framework for playing and learning board games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method uses MCTS only at test time for faster game learning.
Unified framework for N-tuples learning improves weakly supervised tasks.
We develop the concept of a double (more generally n-tuple) principal bundle departing from a compatibility condition for a principal action of a Lie group on a groupoid.
New algorithms for n-player games using a player-centered approach.
The true online TD(λ) algorithm has recently been proposed (van Seijen and Sutton, 2014) as a universal replacement for the popular TD(λ) algorithm, in temporal-difference learning and reinforcement learning. True online TD(λ) has better theoretical properties than conventional TD(λ), and the expectation is that it als…
RO-TD learns sparse value functions efficiently.
Introduces TDRC to balance TD's ease and soundness.
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side. In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms and provide theoretical analysis on their conver…
Improved TD learning reduces batch sampling error.
Uniform TD(0) bound derived for function approximation with Markov noise.
New TD algorithms stabilize RL tasks by reformulating updates into fixed point equations.
New TD method stabilizes average-reward learning.
Improved TD(0) algorithm for reinforcement learning with linear approximations.
In reinforcement learning, the TD() algorithm is a fundamental policy evaluation method with an efficient online implementation that is suitable for large-scale problems. One practical drawback of TD() is its sensitivity to the choice of the step-size. It is an empirically well-known fact that a large step-size l…
Improved TD learning reduces variance and bias errors.
Our understanding of reinforcement learning (RL) has been shaped by theoretical and empirical results that were obtained decades ago using tabular representations and linear function approximators. These results suggest that RL methods that use temporal differencing (TD) are superior to direct Monte Carlo estimation (M…
Improved TD learning with tail averaging and regularization achieves optimal convergence rates.
In this paper, we introduce a method for adapting the step-sizes of temporal difference (TD) learning. The performance of TD methods often depends on well chosen step-sizes, yet few algorithms have been developed for setting the step-size automatically for TD learning. An important limitation of current methods is that…
Paper proposes an active multi-step TD algorithm for reinforcement learning.
The family of temporal difference (TD) methods span a spectrum from computationally frugal linear methods like TD(λ) to data efficient least squares methods. Least square methods make the best use of available data directly computing the TD solution and thus do not require tuning a typically highly sensitive learning r…
The study proposes using TD error for selecting σ in Q(σ, λ).
Study on distributional TD learning with linear approximations for better return estimation.
While there are convergence guarantees for temporal difference (TD) learning when using linear function approximators, the situation for nonlinear models is far less understood, and divergent examples are known. Here we take a first step towards extending theoretical convergence guarantees to TD learning with nonlinear…
We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (TD) learning and Monte Carlo (MC) policy evaluation. The two methods are known to achieve complementary bias-variance trade-off properties, w…
Unified framework for finite-sample RL algorithms using Lyapunov theory.
The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy improvement scheme. However, most Temporal Difference (TD) based solutions ignore the discrepancy between the stationary distribution of the beh…
In this paper, we provide a unified analysis of temporal difference learning algorithms with linear function approximators by exploiting their connections to Markov jump linear systems (MJLS). We tailor the MJLS theory developed in the control community to characterize the exact behaviors of the first and second order …
TD learning reduces interference, leading to better generalization.
Accelerates TD learning for long-horizon reinforcement learning problems.
Temporal-difference (TD) networks are a class of predictive state representations that use well-established TD methods to learn models of partially observable dynamical systems. Previous research with TD networks has dealt only with dynamical systems with finite sets of observations and actions. We present an algorithm…
Paper shows TD learning without projection converges robustly.
Maps from buildings to spaces study K-theory of Hecke algebras.
Paper improves TD(0) convergence rate with LFA, i.i.d. samples, and averaging.
TDS provides exact samples for conditional distributions in diffusion models.
Proposes a convergent TD algorithm for off-policy RL.
Stochastic differential equation approximation for linear TD(0) under Markovian noise
Paper analyzes TD() convergence rates for arbitrary features.
TDprop uses Jacobi preconditioning to improve adaptive optimizers in Deep RL.
Paper presents a technique using Spearman's Rank Correlation Coefficient for KE in TDs.
New algorithms improve distributional TD learning with linear approximations.
In unsupervised domain adaptation (UDA), classifiers for the target domain (TD) are trained with clean labeled data from the source domain (SD) and unlabeled data from TD. However, in the wild, it is difficult to acquire a large amount of perfectly clean labeled data in SD given limited budget. Hence, we consider a new…
TD-Flow improves long-term predictions in agent learning.
TD(0) with Polyak-Ruppert averaging achieves robust and fast convergence rates
We consider two families of algebraic varieties indexed by natural numbers : the configuration space of unordered -tuples of distinct points on , and the space of unordered -tuples of linearly independent lines in . Let be any sequence of virtual -representations give…
Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due to the nonlinearity in value function approximation, such a coupling leads to nonconvexity and even divergence in optimization. As a result, the global convergence …
We present a database of parliamentary debates that contains the complete record of parliamentary speeches from Dáil Éireann, the lower house and principal chamber of the Irish parliament, from 1919 to 2013. In addition, the database contains background information on all TDs (Teachta Dála, members of parliament), such…
We provide non-asymptotic bounds for the well-known temporal difference learning algorithm TD(0) with linear function approximators. These include high-probability bounds as well as bounds in expectation. Our analysis suggests that a step-size inversely proportional to the number of iterations cannot guarantee optimal …