We reduce variance in RL with input-dependent baselines.
problem High variance in RL with standard baselines in input-driven environments.
method Derive and use a bias-free, input-dependent baseline; propose a meta-learning approach.
result Input-dependent baselines improve training stability and policy quality.
New empirical index reveals wider validity of Echo State Property in input-driven reservoirs.
problem Lack of proper input consideration in Echo State Property conditions.
method Introduced an empirical Echo State Property index to analyze stability of reservoirs with input signals.
result The actual domain of Echo State Property validity is wider than literature conditions suggest.
New insights into how encoder-decoder networks generate attention matrices.
problem Understanding how encoder-decoder networks use attention matrices.
method Decomposing hidden states into temporal and input-driven components.
result Attention matrices are formed based on task requirements, not architecture type.
Quantum reservoir computing needs coherence influx for effective information processing.
problem Understanding and optimizing quantum reservoir computing.
method Theoretical and numerical analysis of quantum systems, focusing on coherence influx and spectral radius of Pauli transfer matrix.
result Coherence influx is essential for realizing nonstationary echo state property in quantum reservoir computing.
Unsupervised learning representations generalize better than supervised learning under distribution shifts.
problem Robustness of unsupervised representations to distribution shift.
method Extensive evaluation on synthetic and realistic datasets, including controllable domain generalization datasets.
result Unsupervised representations learned from SSL and AE generalize better than supervised learning under various distribution shifts.
Study evaluates quantum and classical conditional Boltzmann machines for time-series forecasting.
problem Time-series forecasting using quantum and classical conditional Boltzmann machines.
method Developed and compared four conditional energy-based forecasting architectures: Gaussian-Bernoulli CRBM, QCRBM, QQRBM, and QFeatureQRBM. Evaluated using symmetric hyperparameter optimisation.
result No systematic evidence of a quantum advantage in time-series forecasting at the available sample size.
Two algorithms improve Federated RL in diverse environments.
problem Collaborative learning in environments with varying dynamics.
method Proposed two federated RL algorithms, QAvg and PAvg, and a personalization heuristic.
result Achieved better performance and generalization in diverse environments.
OBSER framework infers sub-environments from objects, outperforming scene-based methods.
problem Zero-shot recognition of environments from object distributions.
method Bayesian framework using metric and self-supervised learning models to estimate object distributions in latent space.
result OBSER framework reliably performs inference in open-world and photorealistic environments, outperforming scene-based methods.
UAED discovers adaptive environments for robust learning.
problem Avoiding spurious correlations in data.
method Unified framework that learns a distribution over data transformations.
result Improves worst-case accuracy on standard benchmarks.
New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
problem Challenges in planning for stochastic and partially-observable environments.
method Uses discrete autoencoders and a stochastic variant of Monte Carlo tree search.
result Significantly outperforms MuZero on stochastic chess and scales to DeepMind Lab.
Self-supervised policy adapts after deployment without rewards.
problem Generalizing reinforcement learning policies across different environments.
method Uses self-supervision to train policies in new environments without reward signals.
result Significant improvements in generalization across diverse environments.
Bayesian model for multi-environment prediction with latent variable changes.
problem Prediction in environments with changing latent variable distributions.
method Bayesian model with empirical Bayes prior and amortized variational algorithm.
result Method outperforms previous approaches in new environments.
This paper introduces CENIE to quantify environment novelty for better UED.
problem Challenges in measuring environment novelty for effective UED.
method CENIE framework using state-action space coverage and Gaussian Mixture Models.
result CENIE improves UED performance across multiple benchmarks.
A new method learns to prioritize and use multiple views of an environment for better decision-making.
problem Learning from multiple views of an environment to improve decision-making.
method Attention-based deep reinforcement learning to dynamically attend to views of the environment.
result The method improves performance in complex 3D environments with obstacles.
Proposes a Kalman Filter modifier to improve neural network performance in changing environments.
problem Maintaining performance of neural networks in non-stationary environments.
method Kalman Filter based modifier to adapt to changes.
result The proposed model adapts better to changes with a 0.4% accuracy drop compared to 90% for conventional models.
Infinite hierarchical contrastive clustering identifies personal environments linked to health outcomes.
problem Identifying meaningful relationships between environmental features and health outcomes on an individual level.
method Contrastive clustering framework with stick-breaking prior and participant-specific prediction loss.
result Model effectively identifies distinct personal environments and groups them into meaningful types linked to health outcomes.
We consider apprenticeship learning, i.e., having an agent learn a task by observing an expert demonstrating the task in a partially observable environment when the model of the environment is uncertain. This setting is useful in applications where the explicit modeling of the environment is difficult, such as a dialog…
This paper tackles RL in non-stationary environments, improving decision-making.
problem Develop optimal RL decisions in non-stationary environments.
method Adapted change point algorithm for detecting model changes and developed an RL algorithm.
result RL algorithm maximizes long-run reward in changing environments.
Research shows collective learning across diverse environments is hard due to privacy and security concerns.
problem Privacy, security, and equity concerns restrict information sharing in diverse AI environments.
method Characterized learning algorithms as choice correspondences, provided minimum requirements for rational learning algorithms.
result The only rational learning algorithm in heterogeneous environments is unilaterally learning from a single environment without information sharing.
LEADS improves model generalization across different environments.
problem Modeling dynamical systems from varied environments leads to biased or scarce solutions.
method LEADS learns a shared model capturing common dynamics and additional terms for environment-specific dynamics.
result LEADS improves model generalization for both known and novel environments.
Proposes a method to improve reinforcement learning in multi-view environments.
problem Improving policy learning in environments with partial observability.
method Actor-Critic-Attention Mechanism for generating a single feature representation from multiple views.
result Our method outperforms state-of-the-art baselines on various complex 3D environments.
A new method shapes reinforcement learning environments by abstracting large state spaces.
problem Learning in large, noisy environments with sparse feedback.
method Environment shaping using state abstraction.
result Agent's policy in shaped environment preserves near-optimal behavior in original environment.
TOYBOX creates new Atari environments for deep RL experiments.
problem Challenges in evaluating deep RL agent behavior.
method Designs and opensource releases a subset of Atari environments.
result TOYBOX enables experiments impossible in other environments.
Generative model learns spatial memory for predicting agent's future in complex environments.
problem Training generative and temporal models in partially observed, 3D environments is challenging.
method Action-conditioned generative model with non-parametric spatial memory and state-space model for dynamics.
result Scalable architecture capable of coherent predictions over hundreds of time steps in 2D and 3D environments.
MiniHack simplifies creation of complex RL environments.
problem Limited availability of challenging RL benchmarks.
method Develops a sandbox framework for easy RL environment design.
result MiniHack enables rapid creation of diverse RL testbeds.
WILD-SCAV benchmarks AI in complex 3D FPS environments.
problem Lack of complexity and diversity in RL environments.
method Developed a 3D open-world FPS game environment.
result Demonstrates effectiveness in benchmarking RL algorithms.
Paper proposes DEMER to reconstruct hidden confounders for better reinforcement learning in recommendation.
problem Reinforcement learning in real-world applications is costly due to exploration in the environment.
method DEMER uses a multi-agent generative adversarial imitation learning framework to learn the environment and hidden confounder.
result DEMER effectively reconstructs hidden confounders and improves recommendation policy performance.
Enhances speech recognition in new environments by embedding noise and scaling training data.
problem Improving speech recognition in unseen noisy environments.
method Embedding noise from unseen environments and scaling training data to 16,784 environments.
result Reduced word error rate from 34.04% to 15.46% on enhanced speech.
PSRL extension for continuing environments reduces regret.
problem Formalizing and analyzing resampling approach for reinforcement learning.
method Continuing PSRL maintains a model of the environment and replaces it with samples from the posterior distribution.
result Established an i l d e O ( τ S A T ) ilde{O}(τS \sqrt{A T}) i l d e O ( τ S A T ) bound on Bayesian regret. Study evaluates self-supervised representations in interactive environments.
problem Unclear which self-supervised methods best capture meaningful features of diverse environments.
method Quantitative evaluation of representations in two visual environments: Flappy Bird and Sonic The Hedgehog.
result Representations' utility depends on the visuals and dynamics of the environment.
APES simplifies reinforcement learning environment design in Python.
problem Creating and simulating reinforcement learning environments.
method Introduces APES, a Python toolbox for 2D grid-world environments.
result Equips reinforcement learning agents with customizable field of vision and item/reward interactions.
New method learns robust representations by modeling environment variation.
problem Learning invariant representations across varying environments.
method Explicitly modeling variation across environments and marginalizing it out.
result Proposed method outperforms invariant-learning methods in various settings.
GALA framework learns invariant graph representations via environment augmentation with minimal assumptions.
problem Learning invariant graph representations from different environments without additional assumptions.
method Developed GALA framework with minimal assumptions of variation sufficiency and consistency. Uses an assistant model to differentiate graph environment changes.
result Extracting maximally invariant subgraphs to proxy predictions identifies underlying invariant subgraphs for successful out-of-distribution generalization.
Optimal attack against autoregressive models by manipulating environment states.
problem Manipulating autoregressive forecasts to track a target trajectory.
method Linear Quadratic Regulator (LQR) for linear models, Model Predictive Control (MPC) for nonlinear models.
result Optimal attack formulations for both white-box and black-box settings.
Generative models help agents form stable beliefs in complex environments.
problem Forming and maintaining stable beliefs in complex, dynamic environments.
method Train expressive generative models to predict multiple steps ahead.
result Expressive generative models improve data-efficiency in RL tasks.
ATLAS separates invariant and transferable latent factors across diverse environments.
problem Transfer learning and robust prediction in heterogeneous environments.
method ATLAS leverages invariance principle to disentangle latent factors and uses auxiliary labels for robust prediction.
result Near-oracle performance and robust transferable prediction in new environments.
XRM discovers environments without human annotations for OOD methods.
problem Costly and biased manual annotations limit OOD methods.
method XRM trains twin networks to mimic mistakes, eliminating hyper-parameters.
result XRM achieves oracle worst-group-accuracy for OOD methods.
Advantage amplification helps RL in slow-evolving latent-state environments.
problem Challenges in reinforcement learning for long-horizon latent-state environments.
method Temporal abstraction and aggregation methods to overcome belief state error and small action advantage.
result Proven advantage amplification in settings with slowly evolving latent states.
We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment class in the sense that (1) asymptotically its …
Method learns attractive areas from agent motions to represent environments.
problem Representing environments based on moving agents' nonlinear motions.
method Switching model of velocity fields, parametric representation of attractive spots.
result Dynamic map of attractive areas for online learning.
AR-A3C improves A3C's robustness to noisy environments.
problem Neural networks are vulnerable to noise from unexpected sources in reinforcement learning.
method Introduced an adversarial agent to make the learning process more robust.
result AR-A3C outperforms A3C in both clean and noisy environments.
Symmetry-based learning needs interaction with the environment.
problem Defining disentangled representations in dynamic environments.
method Building on Symmetry-Based Disentangled Representation Learning, we argue that agents need to interact with the environment to discover symmetries.
result Agents need to interact with the environment to fully understand symmetries and learn disentangled representations.
Researchers use DT to transfer policies from one environment to another using causal reasoning.
problem Adapting to changes in environmental dynamics in reinforcement learning.
method Applying causal counterfactual reasoning to Decision Transformer (DT) architecture for policy transfer.
result DT successfully transfers a learned policy to new environments while retaining most of the reward.
Agent learns policies from world models in simulated environments.
problem Training reinforcement learning agents in diverse environments.
method Generative neural network models environments, compact policies evolved from features.
result Achieves state-of-the-art performance across various environments.
New algorithm guarantees domain generalization with few environments.
problem Performing well on unseen environments with limited training data.
method Iterative feature matching algorithm with theoretical guarantees.
result Guaranteed domain generalization with logarithmic environments.
NetHack Learning Environment (NLE) tests RL algorithms, offering scalable, complex, and challenging gameplay.
problem Challenging environments for testing RL algorithms.
method Procedurally generated, stochastic, rich, and complex NetHack environment.
result Demonstrates empirical success for early stages of NetHack using RL.
New method uses autoregressive models for dynamic planning.
problem Planning in dynamic environments with moving obstacles and goals.
method Conditional autoregressive generative models in a discrete latent space.
result Method nearly matches true environment performance for planning.
This paper introduces Dex, a reinforcement learning environment toolkit specialized for training and evaluation of continual learning methods as well as general reinforcement learning problems. We also present the novel continual learning method of incremental learning, where a challenging environment is solved using o…