Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

20416181 · Jun 202019922001200920182026
48 results for Distal Rewards

Convolutional network predicts DNA chromatin structure from sequence images.

problem Predicting chromatin structure from DNA sequences.
method Developed a convolutional neural network using image-representation of DNA sequences.
result The method outperforms existing methods in prediction accuracy and training time.

CCAT improves model robustness to various adversarial attacks.

problem Robustness to adversarial attacks does not generalize to unseen threat models.
method CCAT biases models towards low confidence predictions on adversarial examples.
result CCAT increases robustness against multiple adversarial attack norms and types.

New method improves mHealth user engagement using Thompson sampling for count data.

problem Optimizing mHealth interventions for distal outcomes through proximal context.
method Combines count data models with Thompson sampling for contextual bandits.
result Improves user engagement in mHealth trials compared to existing methods.

A new model explains protein interactions via electron delocalization.

problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.

StepMix estimates mixture models with covariates for social science applications.

problem Estimating latent classes with covariates in social science models.
method Pseudo-likelihood estimation using one-, two-, and three-step approaches.
result Unified framework for expectation-maximization subroutines.

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

Paper addresses reward learning issues in RL, improving both under- and over-estimation.

problem Reward learning from data can lead to reward delusions or underestimation, causing unintended behaviors.
method Connects reward learning to positive-unlabeled (PU) learning and applies a large-scale PU learning algorithm.
result Improves both GAIL and supervised reward learning without additional assumptions.

The study categorizes reward errors in reinforcement learning, finding some can be beneficial.

problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.

Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.

problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

Proposes a method to boost deep reinforcement learning with sparse rewards.

problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.

Extends Hindsight Experience Replay to learn from multiple reward functions.

problem Learning policies for multiple reward functions without trial runs.
method Develops a method to use a reward function to calculate rewards for states not encountered, and learns policies for multiple reward functions.
result A single policy can generalize across all linear combinations of multi-objective rewards.

Action guidance helps agents learn true objectives in games with sparse rewards.

problem Training agents in games with sparse rewards requires significant exploration.
method Action guidance, a novel technique that combines exploration with reward shaping.
result Action guidance enables agents to optimize true objectives efficiently.

New RL method uses distance between states instead of rewards for sparse reward environments.

problem Sparse rewards or non-reward environments in reinforcement learning.
method Uses goal-distance gradient and bridge point planning for policy improvement.
result Significantly better performance on sparse reward and local optimal problems in complex environments.

Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.

problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.

Enhances reward specification in RL with a novel language-based approach.

problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.

Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.

problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.

RUDDER simplifies Q-value estimation for delayed rewards in MDPs.

problem Solving delayed rewards in reinforcement learning with bias and variance issues.
method Reward redistribution and return decomposition to simplify Q-value estimation.
result RUDDER significantly speeds up Q-value estimation and improves performance on Atari games.

Extends reinforcement learning alignment to scalar rewards, improving math reasoning.

problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.

BelMan uses Bayesian methods to optimize decisions in multi-armed bandit problems.

problem Optimizing decisions in multi-armed bandit problems with varying rewards and beliefs.
method BelMan uses a geometric approach with information projection and reverse projection to balance exploration and exploitation.
result BelMan outperforms other algorithms in specific scenarios involving many arms and continuous rewards.

Active Inverse Reward Design improves AI agent training by querying users for reward function preferences.

problem Iterative reward function tuning in AI agents is inefficient and may not generalize well.
method Structured queries to the user to compare reward functions, updating posterior with IRD.
result Substantially outperforms IRD in test environments, inferring non-linear rewards.

This work characterizes reward function partial identifiability and its impact on policy optimization.

problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.

This paper introduces a new reward shaping method for average-reward reinforcement learning.

problem Speeding up convergence to an optimal policy in average-reward reinforcement learning tasks.
method Developed a temporal logic-based approach to automatically generate reward shaping functions.
result The optimal policy can be recovered using the proposed reward shaping framework.

Reinforcement learning struggles with corrupted reward signals.

problem Agents observe rewards that are not accurate due to sensory errors or software bugs.
method Formalized as Corrupt Reward MDP, investigated two approaches: richer data and randomisation.
result Traditional RL methods fail in CRMDPs, necessitating new approaches.

Paper generalizes reward distribution in multi-armed bandits with temporally-partitioned rewards.

problem Handling partial rewards distributed over multiple rounds in multi-armed bandits.
method Introduces Beta-spread property to generalize reward distribution, derives lower bound, and provides TP-UCB-FR-G algorithm.
result Improves regret upper bound for some scenarios using Beta-spread property.

Paper tackles noisy reinforcement learning with perturbed rewards, improving agent performance.

problem Noisy rewards in reinforcement learning scenarios, leading to unreliable model performance.
method Develops a robust RL framework using a confusion matrix to estimate unbiased surrogate rewards.
result Trained policies using estimated surrogate rewards achieve higher expected rewards and faster convergence.

Paper proposes a new algorithm to learn intrinsic rewards for policy-gradient agents.

problem Learning intrinsic rewards for policy-gradient agents is an open problem.
method Derives a novel algorithm for learning intrinsic rewards for policy-gradient based learning agents.
result Improved performance on most domains but not all when using the new algorithm.

The paper tackles reinforcement learning with exogenous variables and rewards.

problem Exogenous state variables and rewards slow reinforcement learning by introducing uncontrolled variation.
method Formalizes exogenous state variables and rewards, decomposes MDP into exogenous and endogenous components, and introduces algorithms to discover these components.
result Optimal policies for the endogenous MDP are also optimal for the original MDP, but the endogenous MDP is easier to solve due to reduced variance.

Study optimal reward schemes for inducing desired player performance in risky contests.

problem Designing optimal rewards to encourage desired performance levels in risky contests.
method Analyzed the optimal reward schemes for inducing average and specific rank performance.
result Optimal reward schemes can have surprising shapes, not just related to inequality.

We formalize and decompose reinforcement learning problems with exogenous state variables and rewards.

problem Exogenous state variables and rewards slow down reinforcement learning.
method Formalized exogenous state variables and rewards, decomposed MDPs, derived variance-covariance condition, developed algorithms.
result Monte Carlo policy evaluation on the endogenous MDP is accelerated compared to using the full MDP.

Bayesian nonparametric models improve multi-armed bandit performance with uncertain rewards.

problem Reward model uncertainty in multi-armed bandits.
method Bayesian nonparametric Gaussian mixture models for flexible reward density estimation.
result Achieves successful regret performance with asymptotic regret bound.

Study scaling laws of reward model overoptimization in reinforcement learning.

problem Reward model overoptimization hinders true performance in reinforcement learning.
method Synthetic setup with fixed gold reward model; optimization using RL or best-of-nn sampling; analysis of scaling laws.
result Scaling laws of reward model overoptimization differ based on optimization method and scale smoothly with model parameters.

Develops HMRL for sparse reward RL problems, improving meta policy efficiency and transferability.

problem Difficulty in learning meta policies for sparse reward RL problems.
method Hyper-Meta RL framework with cross-environment meta state embedding and shaped meta reward.
result Improves meta policy generalization and efficiency for sparse reward RL problems.

Natural language instructions improve reinforcement learning efficiency.

problem Designing effective reward functions for reinforcement learning is difficult and time-consuming.
method Proposes LanguagE-Action Reward Network (LEARN) to map natural language instructions to intermediate rewards.
result Language-based rewards lead to successful task completion 60% more often than without language.

Method predicts future rewards from past actions in a linear Gaussian system.

problem Maximizing cumulative reward in a stochastic multi-armed bandit with linear Gaussian dynamics.
method Proposes a method using a modified Kalman filter to predict future rewards based on past rewards.
result Reward from any action can be used to predict another action's future reward.

Paper addresses reward hacking in preference optimization, proposing POWER-DL to improve AI alignment.

problem Reward hacking problem in preference optimization, leading to undesired behaviors.
method POWER-DL combines robust reward maximization and dynamic label updates to mitigate reward hacking.
result POWER-DL consistently outperforms state-of-the-art methods on alignment benchmarks.