Reward estimation improves model-free RL performance under corrupted rewards.
problem Handling corrupted or stochastic rewards in reinforcement learning.
method Using an estimator for both rewards and value functions.
result Improves performance in various noise types and environments.
Enhances on-policy RL with past reward statistics and hot-wiring.
problem Challenges of sparse reward signals in on-policy methods.
method Multi-critic supervision and hot-wiring mechanism.
result Improves on-policy learning for sparse reward tasks.
Proposes a method to boost deep reinforcement learning with sparse rewards.
problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.
Solves reinforcement learning tasks by guiding policies with a penalty signal.
problem Learning policies that exploit reward loopholes and misspecifications.
method Multi-timescale approach using a penalty signal for constraint satisfaction.
result Proves the convergence of Reward Constrained Policy Optimization (RCPO).
New algorithm learns reward signals from episodic returns for better reinforcement learning.
problem Difficulty in designing reward functions for real-world reinforcement learning tasks.
method Introduces a new algorithm that decomposes episodic returns into time-step rewards using deep neural networks.
result Learning reward signals from episodic returns improves reinforcement learning efficiency.
Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.
problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.
Paper introduces a novel reward function for noisy financial markets using imitation learning.
problem Noisy reward function in financial markets hinders RL agent performance.
method Integrates imitation learning feedback with reinforcement learning to improve reward function design.
result Improves financial performance metrics compared to traditional benchmarks and RL agents.
EMI uses predictive signals to guide exploration in sparse reward settings.
problem Challenges of reinforcement learning with sparse reward signals.
method Constructs embedding representations of states and actions for forward prediction in the representation space.
result Competitive results on challenging tasks with continuous control and discrete actions.
Self-supervised reward prediction improves RL in sparse reward settings.
problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.
A new method shapes reinforcement learning environments by abstracting large state spaces.
problem Learning in large, noisy environments with sparse feedback.
method Environment shaping using state abstraction.
result Agent's policy in shaped environment preserves near-optimal behavior in original environment.
Trade-R1 bridges verifiable rewards to stochastic financial markets via process-level reasoning verification.
problem Extending RL to financial markets where rewards are verifiable but noisy.
method A verification method that transforms reasoning over financial documents into a structured RAG task, using a triangular consistency metric.
result DSR achieves superior cross-market generalization while maintaining reasoning consistency.
A new learning method for prosthetic arms without explicit rewards.
problem Learning a prosthetic arm to interact with users without explicit reward signals.
method Interaction-Grounded Learning, observing multidimensional context and feedback vectors, discovering latent reward signal.
result The algorithm can discover a latent reward signal and ground its policies for successful interaction.
Q-Learning overestimation bias influenced by learning rate, discount factor, and reward signal.
problem Overestimation bias in Q-Learning algorithm.
method Investigated the influence of learning rate, discount factor, and reward signal on Q-Learning's overestimation bias. Tuned parameters and used an exponential moving average of reward signal.
result Q-Learning can achieve more accurate value estimates by tuning parameters and using an exponential moving average of reward signal.
A method using competitive experience replay enhances learning from sparse rewards.
problem Learning from sparse rewards in reinforcement learning.
method Competitive experience replay method that augments sparse rewards through an exploration competition between agents.
result The method leads to faster convergence and improved task performance.
Paper shows RLHF can be solved similarly to standard RL.
problem Difficulty of RLHF compared to standard RL.
method Reduction to reward-based RL techniques.
result RLHF can be solved using existing algorithms for reward-based RL.
Unified approach combines reward maximization and empowerment for RL.
problem Combining reward maximization and empowerment for reinforcement learning.
method Unified Bellman optimality principle for empowered reward maximization.
result Unified approach leads to improved initial and competitive final performance.
Reinforcement learning struggles with corrupted reward signals.
problem Agents observe rewards that are not accurate due to sensory errors or software bugs.
method Formalized as Corrupt Reward MDP, investigated two approaches: richer data and randomisation.
result Traditional RL methods fail in CRMDPs, necessitating new approaches.
We formalize and decompose reinforcement learning problems with exogenous state variables and rewards.
problem Exogenous state variables and rewards slow down reinforcement learning.
method Formalized exogenous state variables and rewards, decomposed MDPs, derived variance-covariance condition, developed algorithms.
result Monte Carlo policy evaluation on the endogenous MDP is accelerated compared to using the full MDP.
Paper addresses reward learning issues in RL, improving both under- and over-estimation.
problem Reward learning from data can lead to reward delusions or underestimation, causing unintended behaviors.
method Connects reward learning to positive-unlabeled (PU) learning and applies a large-scale PU learning algorithm.
result Improves both GAIL and supervised reward learning without additional assumptions.
SEMI uses multisensory incongruity to self-supervise exploration in reinforcement learning.
problem Efficient exploration in reinforcement learning with sparse or missing rewards.
method SEMI incentivizes exploration by maximizing multisensory incongruity, measured in perception and action incongruity.
result SEMI improves sample efficiency and learns skills without external rewards.
The paper examines how updates to probabilistic models influence behavior based on evidence.
problem Understanding how updates to probabilistic models influence behavior based on evidence.
method Study of KL-regularized soft updates as Bayesian posterior updates within a single probabilistic model.
result Posterior updates determine relative incentives but not absolute rewards, which are ambiguous up to context-specific baselines.
Adapts deep reinforcement learning to ordinal rewards.
problem Using numerical rewards in reinforcement learning has drawbacks; ordinal rewards offer an alternative.
method Develops a general approach to converting reinforcement learning algorithms to ordinal reward systems, including Ordinal Deep Q-Networks.
result Ordinal Deep Q-Networks perform comparably to numerical variants on engineered problems and better on simpler reward signals.
Paper proposes a new method for efficient exploration in reinforcement learning.
problem Sparse reward reinforcement learning challenges in exploration.
method Learn separate intrinsic and extrinsic task policies, schedule between them, and use successor feature control (SFC).
result Substantially improved exploration efficiency with SFC and hierarchical usage of intrinsic drives.
EBIL simplifies IL by estimating expert energy as reward, achieving effective performance.
problem Recovering optimal policy from expert demonstrations without reward signals.
method EBIL uses a two-stage solution: first estimating expert energy as reward, then learning policy.
result EBIL achieves effective performance and interpretable reward signals.
HyperX uses reward bonuses to enable efficient exploration in meta-learning.
problem Catastrophic failure of meta-learning with sparse rewards.
method HyperX uses novel reward bonuses to explore in approximate hyper-state space.
result HyperX meta-learns better task-exploration and adapts more successfully to new tasks.
SAC-X enables learning complex behaviors from sparse rewards.
problem Learning complex behaviors from sparse reward signals.
method Scheduled Auxiliary Control (SAC-X) with auxiliary tasks.
result SAC-X enables efficient exploration and complex behavior learning.
DOPL learns from preference feedback to solve RMAB problems.
problem Learning optimal decisions in RMAB with limited reward information.
method Direct online preference learning (DOPL) for Pref-RMAB.
result DOPL achieves sublinear regret for RMAB with preference feedback.
Paper introduces OTR for efficient offline RL in surgical robotics.
problem Lack of annotated datasets for offline RL in surgical robotics.
method OTR algorithm using Optimal Transport to assign rewards to unlabeled trajectories.
result OTR enables efficient policy learning from large datasets without handcrafted rewards.
GENE tackles sparse reward in RL by generating states to explore and exploit.
problem Sparse reward in reinforcement learning.
method Generative Exploration and Exploitation (GENE) method.
result GENE significantly outperforms existing methods in tasks with binary rewards.
This work tackles the challenge of aligning generative models without explicit reward signals.
problem Aligning generative models without explicit reward signals.
method A Bilevel Optimization framework where the reward function is treated as the optimization variable of an outer-level problem.
result Theoretical analysis and insights generalize to tabular classification and model-based reinforcement learning.
Paper uses IRL to generate more diverse text.
problem Reward sparsity and mode collapse in text generation.
method Employ inverse reinforcement learning to learn a reward function and optimize policy alternately.
result Our method generates higher quality texts with more diversity.
Meta-learning curiosity algorithms improves exploration across various tasks.
problem Generating curious behavior in reinforcement learning.
method Meta-learning approach to adapt reward signals dynamically.
result Two novel curiosity algorithms outperform human-designed ones.
The paper tackles reinforcement learning with exogenous variables and rewards.
problem Exogenous state variables and rewards slow reinforcement learning by introducing uncontrolled variation.
method Formalizes exogenous state variables and rewards, decomposes MDP into exogenous and endogenous components, and introduces algorithms to discover these components.
result Optimal policies for the endogenous MDP are also optimal for the original MDP, but the endogenous MDP is easier to solve due to reduced variance.
Study shows curiosity-driven learning can perform well without extrinsic rewards.
problem Lack of scalable methods for intrinsic reward design in reinforcement learning.
method Performed a large-scale study of curiosity-driven learning across 54 environments, using prediction error as reward.
result Curiosity-driven learning can achieve good performance without extrinsic rewards, aligning with hand-designed rewards in many cases.
New algorithms for multivariate RL improve decision-making in complex systems.
problem Complex multi-objective decision-making in reinforcement learning.
method Oracle-free and computationally-tractable algorithms for multivariate distributional RL.
result Convergence rates match scalar reward settings and provide insights into reward dimensionality.
Co-DQL improves traffic signal control using multi-agent reinforcement learning.
problem Optimizing signal timing for large-scale traffic control.
method Cooperative double Q-learning (Co-DQL) with mean field approximation and reward allocation.
result Co-DQL reduces average waiting time for vehicles in the road system.
Extends linear MDP to handle nonlinear rewards.
problem Restrictive linear MDP assumption limits real-world applicability.
method Proposes Generalized Linear MDP (GLMDP) with GLMs for rewards.
result Develops offline RL algorithms achieving suboptimality guarantees.
Self-distillation improves constrained language generation by aligning models with target distributions.
problem Sparse and uninformative reward signals in constrained generation settings.
method Iteratively refining the base model through self-distillation, incorporating learned twist functions and proposals.
result Substantial gains in generation quality through improved model alignment with target distributions.
New approach uses unlabeled prior data to accelerate exploration in sparse reward tasks.
problem Sparse reward tasks in reinforcement learning.
method Learn reward model from online experience, label prior data, and use concurrently.
result Rapid exploration in challenging sparse-reward domains.
New policy reduces spectrum access regret in uncoordinated systems.
problem Uncoordinated spectrum access with user-dependent rewards.
method Multi-user multi-armed bandit (MAB) model with user-specific rewards.
result Achieves O(logT) regret for spectrum access. Boosted GFlowNets improve exploration by sequentially training GFlowNets with residual rewards.
problem GFlowNets struggle to evenly explore reward landscapes, leading to poor coverage of high-reward areas.
method Sequential training of an ensemble of GFlowNets, each optimizing a residual reward.
result Boosted GFlowNets achieve better exploration and sample diversity on multimodal benchmarks and peptide design tasks.
BERT embeddings improve sequence quality metrics.
problem Measuring the quality of generated sequences against references.
method Employ contextual BERT embeddings for sequence-level reward.
result Contextual embeddings provide a more effective learning signal.
Graph Denoising Policy Network learns robust representations from noisy graphs.
problem Noise sensitivity in graph representation learning.
method Reinforcement learning to select signal neighborhoods and aggregate features.
result Significantly outperforms state-of-the-art methods on node classification tasks.
A new framework estimates expert policy support to create a reward function for imitation learning.
problem Imitation learning from expert trajectories without reinforcement signals.
method Estimating the support of the expert policy to compute a fixed reward function.
result Comparable or better performance than state-of-the-art methods on discrete and continuous domains.
A new reward learning module improves imitation learning in high-dimensional environments.
problem Challenges in high-dimensional environments for imitation learning.
method Generative model to generate intrinsic reward signals.
result Our method outperforms state-of-the-art IRL methods on Atari games.
TRAIL improves robot imitation learning by focusing on task-relevant features.
problem Discriminator networks learn spurious associations, providing poor reward signals.
method Constrained discriminator optimization to learn task-relevant rewards.
result TRAIL outperforms GAIL and behaviour cloning in robotic manipulation tasks.
A new reinforcement learning method uses mutual information to encourage agents to control their environment.
problem Learning from internal drives instead of external rewards.
method Formulate an intrinsic objective as mutual information between goal states and controllable states, derive a surrogate objective for efficient optimization.
result Demonstrated the efficacy of the approach in robotic tasks.