The paper analyzes an actor-critic algorithm with target networks for deep reinforcement learning.
problem Lack of theoretical understanding of target networks in actor-critic methods.
method Proposes a theoretical analysis of an online target-based actor-critic algorithm with linear function approximation.
result Establishes asymptotic convergence results and finite-time analysis for both critic and actor.
Neural networks trained with actor-critic algorithms converge to ODEs under weak convergence analysis.
problem Challenges in convergence analysis due to changing data distributions in online learning.
method Geometric ergodicity of data samples, Poisson equation, weak convergence techniques.
result Actor and critic networks converge to solutions of ODEs with random initial conditions.
Hybrid actor-critic learns in complex action spaces.
problem Learning in complex, structured action spaces.
method Parallel sub-actor networks and a critic network.
result Hybrid PPO outperforms previous methods in parameterized action spaces.
Model visualizes and analyzes multilayer networks in a latent space.
problem Characterize multiple social networks over a common set of actors.
method Bayesian statistical model with hierarchical prior distribution.
result Visualizes multilayer network data in a low-dimensional Euclidean space.
In a dynamic social or biological environment, the interactions between the actors can undergo large and systematic changes. In this paper we propose a model-based approach to analyze what we will refer to as the dynamic tomography of such time-evolving networks. Our approach offers an intuitive but powerful tool to in…
Proposes online learning for Hawkes processes with network structure and event interaction.
problem Modeling complex interactions and latent structures in network events.
method Online learning approach for mixture of multivariate Hawkes processes.
result Efficacy demonstrated on synthetic and real-world data.
Data mining reveals power structures in Bangladeshi newspapers.
problem Understanding the power dynamics and narrative structure in news reporting.
method Named entity recognition to create temporal actor networks from news statements.
result Cliquishness among powerful political leaders in news articles.
Single-timescale actor-critic finds globally optimal policy.
problem Finding globally optimal policy in reinforcement learning.
method Simultaneous actor and critic updates with linear or deep neural network approximations.
result Actor sequence converges to globally optimal policy at O(K−1/2) rate. Enhances exploration in hierarchical networks using mutual information.
problem Limited exploration in hierarchical Deep Q-Networks.
method Adversarial Soft Actor-Critic with mutual information optimization.
result Improves hierarchical network exploration through mutual information maximization.
Paper analyzes NAC with neural networks for efficient policy optimization.
problem Improving sample and iteration complexity in policy optimization.
method Entropy regularization, averaging, neural network approximation, and optimization techniques.
result Entropy regularization and averaging ensure stability and sharp sample complexity bounds.
New algorithm solves mean-field control problems using actor-critic learning with moment neural networks.
problem Solving mean-field control problems in continuous time reinforcement learning.
method Gradient-based policy and value function learning with moment neural networks on the Wasserstein space.
result Effective solution for diverse mean-field control problems, including multi-dimensional and nonlinear settings.
Traffic actors' future motion predicted using a hybrid graph model.
problem Predicting long-term behaviors of traffic actors in complex scenes.
method A hybrid graph model with nodes for actors and traffic elements, and edges for interaction types.
result TrafficGraphNet achieves state-of-the-art trajectory prediction accuracy.
This work analyzes how neural networks learn representations in actor-critic algorithms.
problem Theoretical support for neural AC algorithms is limited to linear function approximations.
method Mean-field analysis of a two-timescale learning AC algorithm with overparameterized networks.
result Neural AC finds the globally optimal policy at a sublinear rate in the continuous-time and infinite-width limiting regime.
Smaller actor-critic models lead to performance degradation and overfitting, highlighting the critic's role in value underestimation.
problem Performance degradation and overfitting in actor-critic models with smaller actors.
method Broad empirical investigations and analyses of asymmetric actor-critic setups, exploring techniques to mitigate value underestimation.
result Value underestimation is a key cause of performance degradation in smaller actor-critic models, and the critic plays a crucial role in mitigating this.
ASAC uses actor-critic models to optimize observation selection in medical settings.
problem Optimizing observation selection in costly sequential observation scenarios.
method ASAC framework with selector and predictor networks, using actor-critic models for training.
result ASAC significantly outperforms state-of-the-art methods in real-world medical datasets.
Both generative adversarial networks (GAN) in unsupervised learning and actor-critic methods in reinforcement learning (RL) have gained a reputation for being difficult to optimize. Practitioners in both fields have amassed a large number of strategies to mitigate these instabilities and improve training. Here we show …
Deep actor-critic learning optimizes power control in mobile networks.
problem Optimizing power control in large-scale wireless mobile networks.
method Multi-agent deep reinforcement learning with deep deterministic policy gradient.
result The algorithm maximizes a global utility function in a distributed manner.
Neural policy gradient methods converge globally and sublinearly.
problem Global optimality and convergence of neural policy gradient methods.
method Actor-critic schemes with neural networks, proving global optimality and sublinear convergence rates.
result Neural natural and vanilla policy gradient methods converge to globally optimal policies and stationary points.
GRAC improves reinforcement learning by self-guiding and self-regularizing.
problem Learning divergence and slow updates in reinforcement learning algorithms.
method Self-regularized TD-learning and self-guided policy improvement.
result Achieved or outperformed state-of-the-art results on OpenAI gym tasks.
A new method improves actor-critic RL by integrating HMC, enhancing policy distribution and exploration.
problem Actor-critic RL yields suboptimal policies due to amortization gap and insufficient exploration.
method Integrating Hamiltonian Monte Carlo (HMC) into the actor-critic RL framework.
result Improves policy distribution and exploration, leading to better policy estimates and higher returns.
Proposes a new model for predicting future motion of road actors in autonomous vehicles.
problem Forecasting the long-term future motion of road actors for safe autonomous driving.
method Recurrent graph-based attentional approach with interpretable geometric and social relationships.
result Can produce diverse predictions conditioned on hypothetical or 'what-if' scenarios.
Neural networks solve high-dimensional HJB PDEs with asymptotic guarantees.
problem Solving high-dimensional Hamilton-Jacobi-Bellman PDEs in stochastic control theory.
method Actor-critic machine learning algorithm with a structured critic and biased gradient actor.
result The training dynamics converge to an ODE, ensuring solutions to the original problem.
Framework detects influential actors in disinformation networks.
problem Identifying and countering hostile influence operations on social media.
method Combines NLP, ML, graph analytics, and causal inference.
result 96% precision, 79% recall, 96% PR-curve area for IO detection.
Paper introduces a meta-critic for accelerating off-policy actor-critic learning.
problem Improving sample efficiency in continuous control tasks.
method Meta-critic that meta-learns an additional loss for the actor.
result Online meta-critic learning leads to improved performance in various continuous control environments.
Automatically finds strong neural network topologies for continuous control tasks.
problem Handcrafted neural network architectures limit the performance of Deep Reinforcement Learning.
method Combines Neuroevolution with off-policy training and proposes a novel architecture mutation operator.
result The proposed Actor-Critic Neuroevolution algorithm often outperforms strong baseline methods.
Proposes three decentralized multi-agent reinforcement learning algorithms to reduce network congestion.
problem Finding a joint policy maximizing long-term return in a decentralized multi-agent system.
method Three fully decentralized multi-agent natural actor-critic (MAN) algorithms using linear function approximations.
result The proposed algorithms achieve performance comparable to or better than standard methods in reducing network congestion.
PEARL uses reinforcement learning to improve matrix preconditioners.
problem Learning effective preconditioners for iterative solvers is challenging.
method PEARL employs an actor-critic reinforcement learning framework to learn preconditioners dynamically.
result PEARL outperforms traditional and neural preconditioners in flexibility and solving speed.
New algorithm for multi-agent reinforcement learning with reduced communication.
problem Cooperative learning among multiple agents with limited communication.
method Randomized multi-agent actor-critic algorithm for directed graphs.
result Algorithm solves problem for strongly connected graphs with reduced communication.
Physics-guided reinforcement learning optimizes swimming in turbulent flows.
problem Optimizing swimming efforts to maintain proximity in turbulent environments.
method Physics-informed actor-physicist reinforcement learning algorithm.
result Physics-informed reinforcement learning outperforms standard methods in turbulent flow control.
Improves off-policy RL stability with RIS.
problem Stability issues in off-policy RL due to distributional mismatch.
method Relative Importance Sampling (RIS) for off-policy actor-critic.
result RIS stabilizes RL learning by reducing variance.
Paper develops a new multi-agent reinforcement learning algorithm.
problem Improving policies in a network of communicating agents.
method Develops a multi-agent off-policy actor-critic algorithm using emphatic temporal difference learning.
result Proves convergence of the algorithm under linear function approximation.
Proposes a deep reinforcement learning framework for dynamic multichannel access.
problem Efficient use of limited spectral resources in dynamic multichannel access.
method Deep actor-critic reinforcement learning framework for both single-user and multi-user scenarios.
result Demonstrates improved performance and adaptive ability compared to existing methods.
Graph attacks can be successful with just a few bad nodes.
problem Adversarial attacks on graph neural networks.
method Identifying and exploiting anchor nodes to compromise graph models.
result A few bad nodes can significantly degrade graph model performance.
Proposes a model for multi-agent reinforcement learning with hierarchical graph attention network.
problem Limited transferability of trained policies to new multi-agent tasks.
method Uses hierarchical graph attention network for representation learning and multi-agent actor-critic for policy learning.
result Demonstrates superior performance in mixed cooperative and competitive tasks compared to existing methods.
New framework for network regression models accounting for community structure.
problem Inaccurate modeling of residual dependencies in network regression models.
method Modeling errors as community-based and exploiting exchangeability properties.
result Parsimonious standard errors for regression parameters.
Deep learning model solves high-dimensional PDEs using Actor-Critic approach.
problem Solving high-dimensional nonlinear PDEs efficiently.
method Reformulated PDE into BSDE system, inspired by Actor-Critic algorithm for deep RL.
result Improved model with fewer parameters, faster convergence, and less hyperparameter tuning.
In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies. We show that this problem persists in an actor-critic setting and propose novel mechanisms to minimize its effects on both the actor and the cr…
Improves stability of actor-critic methods by penalizing TD error.
problem Instability in actor-critic methods during learning.
method Regularize actor's learning objective by penalizing critic's TD error.
result Improves stability and overall performance of actor-critic methods.
Method extracts taint flows to classify Bitcoin mining pools.
problem Understanding pseudonymous Bitcoin actors and their transactions.
method Taint analysis and graph embedding methods applied to taint flows.
result Taint flows from the same period show high similarity.
Actor-critic methods solve reinforcement learning problems by updating a parameterized policy known as an actor in a direction that increases an estimate of the expected return known as a critic. However, existing actor-critic methods only use values or gradients of the critic to update the policy parameter. In this pa…
DeepPocket uses graph convolutional reinforcement learning for better financial portfolio management.
problem Maximizing return on investment while managing risk in correlated financial assets.
method Graph convolutional reinforcement learning framework with feature extraction, local information collection, and actor-critic reinforcement learning.
result DeepPocket outperformed market indexes on five real-life datasets over three investment periods, including during the Covid-19 crisis.
Actor-critic converges globally in LQR with ergodic cost.
problem Theoretical understanding of actor-critic algorithm's global convergence.
method Nonasymptotic convergence analysis of actor-critic in linear quadratic regulator (LQR) setting.
result Actor-critic finds globally optimal policy and value function at a linear rate.
New PAC-Bayesian approach stabilizes actor-critic learning.
problem Training instability in actor-critic algorithms.
method Employing PAC-Bayesian bound as the critic training objective.
result Significant improvement in online learning performance.
Deep CNNs predict multiple actor trajectories for safer autonomous driving.
problem Uncertainty and variety of traffic behaviors in autonomous driving.
method Convert actor surroundings into images, feed into deep CNNs.
result Successfully tested on SDVs in closed-course tests.
Actors in realistic social networks play not one but a number of diverse roles depending on whom they interact with, and a large number of such role-specific interactions collectively determine social communities and their organizations. Methods for analyzing social networks should capture these multi-faceted role-spec…
Neural network approximates Bayesian decision-making parameters.
problem Analytical intractability of Bayesian decision-making in naturalistic tasks.
method Neural amortization of Bayesian actor model.
result Efficient gradient-based inference of Bayesian actor model parameters.
AC-RNN improves RNN for sequence labeling tasks.
problem RNN's exposure bias in maximum-likelihood training.
method Actor-Critic training for RNNs.
result AC-RNN outperforms CRF on NER and CCG tagging.
A new policy improvement method using CEM for Actor-Critic.
problem Improving policy efficiency and robustness in reinforcement learning.
method Greedy Actor-Critic (Greedy AC) using Conditional Cross-Entropy Method (CCEM).
result Greedy AC outperforms Soft Actor-Critic and is less sensitive to entropy regularization.