Temporal difference learning explained through gradient splitting, improving convergence times.
problem Learning value functions in Markov Decision Processes with linear approximations.
method Interpreting TD learning as gradient splitting and applying convergence proofs from gradient descent.
result Improved convergence times for TD learning, especially with a minor variation.
Study of convergence in Lorentzian spacetimes using temporal functions.
problem Non-compactness of spacetime isometries and convergence in semi-Riemannian settings.
method Introduced anchored convergence and used Cauchy temporal functions to define convergence for spacetimes.
result Established local and global regularity of Cauchy temporal functions and their properties.
Temporal-difference and Q-learning learn feature representations that converge to optimal ones.
problem Understanding how feature representations evolve in temporal-difference and Q-learning with neural networks.
method Mean-field theory applied to overparameterized two-layer neural networks.
result The feature representation converges to the optimal one, generalizing previous results.
Study sharp convergence rates of empirical UOT for spatio-temporal point processes.
problem Statistical analysis of UOT for spatio-temporal point processes.
method Empirical plug-in estimators for Kantorovich-Rubinstein distance between intensity measures.
result Sharp convergence rates of empirical UOT in terms of intrinsic dimensions of measures.
Neural TD and Q-learning prove to converge globally to optimal solutions.
problem Nonconvexity and divergence in neural TD due to value function approximation.
method Proving global convergence of neural TD and Q-learning using overparametrization of neural networks.
result Neural TD and Q-learning converge globally to the global optimum of mean-squared projected Bellman error.
Quantile Temporal-Difference learning proved convergent with proof.
problem Lack of theoretical understanding of QTD despite empirical success.
method Proof of convergence using stochastic approximation and non-smooth analysis.
result QTD converges to fixed points with probability 1.
A new metric assesses spatio-temporal forecast quality using Gini regularized Optimal Transport.
problem Evaluating the quality of spatio-temporal forecasts.
method Gini-regularized Optimal Transport (OT) problem, using Gini impurity function as a regularizer.
result The Gini regularized OT problem converges to the classical OT problem, offering a numerically more stable algorithm.
Proposes a convergent TD algorithm for off-policy RL.
problem Learning value function from different policies in RL.
method Convergent on-policy TD algorithm with linear function approximation.
result Proposes a convergent TD algorithm for off-policy RL.
Study compares three splitting methods for American option valuation.
problem Valuation of American options using numerical methods.
method Three splitting methods: explicit payoff, Ikonen-Toivanen, Peaceman-Rachford.
result Temporal accuracy of splitting methods compared to penalty approach.
Accelerates TD learning for long-horizon reinforcement learning problems.
problem Slow convergence of conventional TD learning in long-horizon tasks.
method Introduces PID Accelerated Temporal Difference (PID TD) learning algorithms.
result Accelerates convergence of TD learning compared to conventional methods.
COF-PAC converges with novel critic and learning method.
problem Convergent off-policy actor-critic with function approximation.
method Two-timescale approach with Gradient Emphasis Learning (GEM).
result First provably convergent COF-PAC with linear critics and nonlinear actor.
Develops a new framework for temporal anchoring in deep embedding spaces.
problem Temporal anchoring in deep embedding spaces, especially drift and convergence issues.
method Operator-theoretic framework with drift maps and event-indexed blocks, proving convergence theorems and equivalence theorems.
result Proves convergence theorems and equivalence theorems for the proposed framework.
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
problem Eigenvalue distribution of Wishart matrix with temporal correlation.
method Analysis of moments and convergence to deformed Marchenko-Pastur distribution for Gaussian process with temporal correlation.
result Eigenvalue distribution converges to deformed Marchenko-Pastur distribution with longer tail and higher peak.
The paper characterizes conditions for convergence of RL methods with linear approximations.
problem Characterizing non-uniqueness issues for reinforcement learning algorithms with linear function approximation.
method Proves a condition on features that determines convergence or non-uniqueness of natural algorithms.
result Natural algorithms converge to the correct solution if and only if value functions in the approximation space satisfy a certain shape.
Paper analyzes TD( λ λ λ ) convergence rates for arbitrary features.
problem Convergence rates for linear TD( λ λ λ ) under arbitrary features. method Developed a novel stochastic approximation result for arbitrary features.
result Established L 2 L^2 L 2 convergence rates for linear TD( λ λ λ ) without linearly independent features assumption. AdaShift solves non-convergence issue of Adam by decorrelating gradient and second-moment terms.
problem Non-convergence of adaptive learning rate methods like Adam.
method AdaShift decorrelates gradient and second-moment terms by temporal shifting.
result AdaShift solves non-convergence problem of Adam and maintains competitive performance.
This paper formalizes Q Q Q -learning and linear TD convergence using Lean 4.
problem Formalizing convergence properties of Q Q Q -learning and linear TD learning. method Formal verification using Lean 4 theorem prover and Mathlib library.
result Formalized almost sure convergence of Q Q Q -learning and linear TD learning. New target-based TD learning algorithms improve deep Q-learning convergence.
problem Improving convergence of deep Q-learning algorithms.
method Introducing averaging TD, double TD, and periodic TD algorithms.
result Established asymptotic convergence analyses for averaging TD and double TD, and finite sample analysis for periodic TD.
Improved TD learning with tail averaging and regularization achieves optimal convergence rates.
problem Convergence analysis of TD learning with linear function approximation.
method Tail-averaging and regularization applied to TD learning algorithm.
result Achieves optimal O ( 1 / t ) O(1/t) O ( 1/ t ) convergence rate in expectation and with high probability. Federated learning interprets temporal dynamics across clients with graph attention.
problem Interpreting temporal patterns across decentralized, heterogeneous systems with nonlinear dynamics.
method Graph Attention Network for learning state transition models over latent states communicated between clients.
result First interpretable characterization of cross-client temporal interdependencies in decentralized nonlinear systems.
Improved RBFNN for nonlinear system identification.
problem Nonlinear system identification problem.
method Spatio-temporal extension of RBFNN with time-space orthogonality.
result Spatio-temporal RBFNN outperforms standard RBFNN in convergence and error reduction.
A new LSTM model reduces state updates and improves convergence for long sequences.
problem Vanishing gradient problem in RNNs and slow convergence on long sequences.
method Gaussian-gated LSTM (g-LSTM) with a time gate to control neuron updates.
result The g-LSTM model reduces state updates and computes by at least 10x compared to an equivalent LSTM.
New robust TD learning method for critical domains without observing rare events.
problem Learning robust policies in critical domains with rare events.
method Introduces a κ κ κ -operator for robust TD learning, proving convergence and demonstrating superior performance. result Empirical evaluations show superior performance and robustness to small model errors.
Improved TD learning reduces variance and bias errors.
problem Inefficient optimization variance in TD learning.
method Proposed a mathematically solid analysis of VRTD, showing linear convergence rate and reduced variance and bias errors.
result VRTD converges to a fixed-point solution with reduced variance and bias errors compared to vanilla TD.
We present for the first time an asymptotic convergence analysis of two time-scale stochastic approximation driven by `controlled' Markov noise. In particular, both the faster and slower recursions have non-additive controlled Markov noise components in addition to martingale difference noise. We analyze the asymptotic…
Study on eigenvalue distribution of correlated time series deforming the semi-circle law.
problem Eigenvalue distribution of correlated time series differs from the semi-circle law.
method Analysis of Wigner random matrix with temporal correlation.
result Eigenvalue distribution converges to a deformed semi-circle law with longer tail and higher peak.
Unified framework for PE and TD methods in continuous time and space.
problem Policy evaluation and TD learning in continuous settings.
method Martingale characterization for designing PE algorithms.
result Convergent time-discretized algorithms converge to continuous-time counterparts.
SGDm with fixed step-size diverges under covariate shift, similar to a parametric oscillator.
problem SGDm with fixed step-size diverges under covariate shift.
method Approximated learning system as a time-varying system of ODEs and characterized divergence/convergence modes.
result SGDm with fixed step-size can diverge under covariate shift, similar to resonance in oscillators.
Improved TD learning for non-i.i.d. Markovian data.
problem Convergence analysis of two time-scale TD learning under Markovian samples.
method Non-asymptotic convergence analysis of two time-scale TD with gradient correction under Markovian data.
result Two time-scale TD can converge as fast as O(log t/(t^(2/3))) under diminishing stepsize.
Unified analysis of TD learning using MJLS theory for linear function approximators.
problem Characterizing the exact behaviors of TD learning algorithms with linear function approximators.
method Exploiting connections to Markov jump linear systems (MJLS) theory to analyze TD learning algorithms.
result Closed-form expressions for mean and covariance matrix of TD estimation error at any time step.
Paper introduces a new reinforcement learning method with improved performance.
problem Designing and analyzing efficient reinforcement learning algorithms.
method Proximal gradient temporal difference learning (GTD) with accelerated algorithm GTD2-MP.
result GTD algorithms have linear complexity and improved convergence rate.
We consider the estimation of large covariance and precision matrices from high-dimensional sub-Gaussian or heavier-tailed observations with slowly decaying temporal dependence. The temporal dependence is allowed to be long-range so with longer memory than those considered in the current literature. We show that severa…
Paper analyzes convergence of Adam-type RL algorithms under Markovian sampling.
problem Theoretical convergence analysis of Adam-type RL algorithms.
method Develops techniques for analyzing convergence under Markovian sampling.
result PG-AMSGrad and TD-AMSGrad converge to stationary points or global optima at specified rates.
Paper shows TD learning without projection converges robustly.
problem Investigate convergence of TD learning with linear approx.
method Simple unprojected TD(0) with novel self-bounding property.
result TD(0) converges with rate O ~ ( 1 / T ) \widetilde{\mathcal{O}}(1/\sqrt{T}) O ( 1/ T ) . Method improves clarity in forecasting spatio-temporal data.
problem Forecasting spatio-temporal data with clarity and interpretability.
method Supervised semi-nonnegative matrix factorization with frequency regularization.
result Method offers clearer interpretability in forecasting spatio-temporal data.
Novel Bayesian framework for spatio-temporal neuroimaging data.
problem Inference on multi-task sparse hierarchical regression models with complex spatio-temporal dynamics.
method Flexible hierarchical Bayesian framework with Kronecker product covariance structure, majorization-minimization optimization, and Riemannian geometry.
result Improved performance on synthetic and real M/EEG data.
New model infers causal relationships from spatio-temporal data, even with unobserved confounders.
problem Challenges in inferring causal relationships from spatio-temporal data due to unobserved confounders.
method Spatio-Temporal Hierarchical Causal Models (ST-HCMs) that extend hierarchical causal modeling to the spatio-temporal domain, using the Spatio-Temporal Collapse Theorem.
result Validated the effectiveness of ST-HCMs on both synthetic and real-world datasets, demonstrating robust causal inference in complex dynamic systems.
This study improves convergence of two-timescale SA under Markovian noise in reinforcement learning.
problem Stability and convergence of two-timescale stochastic approximations under Markovian noise.
method Introduced a new control strategy for the fast timescale parameter.
result Established almost sure convergence of TDC with eligibility traces under off-policy learning with linear function approximation.
The paper explores how the probability of default estimation changes with temporal correlation decay.
problem Difficulty in estimating the probability of default due to correlations between borrowers.
method Hierarchical Bayesian estimation using beta binomial distribution with temporal correlation.
result A phase transition occurs in the PD estimator, with convergence depending on the power decay index of temporal correlation.
The study uses the Merton model to estimate PD and finds a phase transition affecting convergence speed.
problem Estimating the probability of default (PD) using limited historical data.
method Adopted the Merton model and analyzed phase transitions in default correlation.
result PD estimation converges slowly when temporal correlation decays by power law less than one.
NADPEx uses dropout to enable temporally consistent exploration in reinforcement learning.
problem Achieving temporally consistent exploration in reinforcement learning agents.
method Integrates dropout into reinforcement learning policies to ensure temporal consistency.
result NADPEx outperforms naive exploration and parameter noise in tasks with sparse rewards.
Dynamic sample pruning speeds up spatio-temporal forecasting models.
problem Training deep learning models on large, redundant datasets is computationally expensive.
method Dynamic sample pruning based on real-time learning state.
result Significant acceleration of training speed with improved performance.
Lazy training and mean field regimes studied for TD learning with nonlinear function approximation.
problem Approximating value function for MRP with TD learning and nonlinear functions.
method Lazy training and mean field scaling of parameters analyzed for convergence.
result Lazy training leads to exponential convergence to local/global minimizers, while mean field scaling results in all fixed points being minimizers.
New algorithms improve reinforcement learning for complex tasks.
problem Improving reinforcement learning for complex tasks.
method Developed new algorithms using distributional reinforcement learning and Cram{é}r distance.
result Proved asymptotic almost-sure convergence of new algorithms for neural networks.
pFedGame uses game theory for decentralized federated learning in dynamic networks.
problem Performance bottlenecks, data bias, model convergence issues, and model poisoning attacks in federated learning.
method pFedGame employs game theory to decentralize federated learning, avoiding a central aggregation server and addressing dynamic network challenges.
result pFedGame achieves higher accuracy (over 70%) in heterogeneous data compared to existing methods.
Proposes a scalable algorithm for large-scale probabilistic tensor analysis.
problem Leveraging time constraints to capture evolving tensor data.
method Introduces a new tensor data split strategy and an efficient algorithm for stochastic Alternating Direction Method of Multipliers.
result Demonstrates that P 2 ^2 2 T 2 ^2 2 F is a highly effective and efficiently scalable algorithm. Paper develops Dense NN models for temporal-spatial data with improved performance.
problem Improving predictive performance and robustness in temporal-spatial modeling.
method Fully connected neural networks with ReLU activation, non-asymptotic bounds, manifold modeling, short-range dependence.
result Demonstrates superior performance in temporal-spatial modeling across various synthetic functions.
Decentralized TD learning converges linearly with linear function approximation.
problem Policy evaluation in fully decentralized multi-agent reinforcement learning.
method Temporal-difference learning with linear function approximation, analyzing i.i.d. and Markovian samples.
result Local estimates converge linearly to the optimum under both i.i.d. and Markovian samples.