Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

9.5%19.0%28.4%37.9% · Jun 201919922001200920182026
48 results for High Reward States

Backtracking model predicts state-action pairs leading to high-reward states for efficient RL.

problem Efficiently learning from environments where only a few states yield high reward.
method Backtracking model that predicts state-action pairs leading to high-reward states.
result Improves sample efficiency of RL algorithms across various environments and tasks.

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

Algorithm estimates human decision-making in high-dimensional states with finite-time guarantees.

problem Estimating optimal policies and measures of fit in dynamic decision models with high-dimensional state spaces.
method Single-loop estimation algorithm with stochastic gradient steps for reward maximization.
result Algorithm converges to a stationary solution with finite-time guarantees and approximates maximum likelihood sublinearly.

Proposes learning latent reward model for planning from rewards.

problem Planning in high-dimensional state spaces with limited reward information.
method Directly learns a latent dynamics model from rewards, planning in latent state-space.
result Successfully learns accurate latent reward prediction model, achieving strong performance and high sample efficiency.

Algorithm learns diffusion processes with high-dimensional state spaces.

problem Stochastic control of unbounded diffusion processes with high-dimensional state spaces.
method Adaptive partitioning and learning algorithm that refines discretization based on estimation bias and statistical confidence.
result Established regret bounds that depend on problem parameters, extending to unbounded diffusion processes.

HashReward improves imitation learning in high-dimensional environments by balancing reward generation and dimensionality reduction.

problem Making policies generalize well in high-dimensional state-action spaces, especially in game playing with raw pixel inputs.
method HashReward uses supervised hashing to balance reward generation and dimensionality reduction.
result HashReward outperforms state-of-the-art methods in high-dimensional environments.

ACE improves GFlowNet exploration efficiency by balancing complementary search strategies.

problem Efficient exploration of diverse high-probability regions in GFlowNets.
method Adaptive Complementary Exploration (ACE) trains a separate GFlowNet to search underexplored regions.
result Significantly improves approximation accuracy and diverse state discovery.

New RL approach uses future state and action visitation measures for better exploration.

problem Improving exploration in reinforcement learning.
method Intrinsic reward based on future state and action visitation measures, using contraction operators.
result Policies achieve good state-action space coverage and high performance.

New method uses reward prediction error for efficient exploration.

problem Efficient exploration in reinforcement learning, especially in complex environments.
method Reward prediction error (RPE) as an intrinsic motivation for exploration, combined with a deep reinforcement learning method (QXplore).
result QXplore outperforms state-novelty methods in diverse tasks, especially when state novelty is not correlated with improved reward.

Bayesian REX learns Atari games from demonstrations efficiently.

problem Bayesian reward learning for complex control problems is computationally intractable.
method Bayesian Reward Extrapolation (Bayesian REX) pre-trains a low-dimensional feature encoding and uses preferences to perform fast Bayesian inference.
result Bayesian REX learns Atari games from demonstrations in 5 minutes, competitive with state-of-the-art methods.

New algorithm improves GAIL for image sequences with global encoder and reward penalization.

problem Low-level, high-dimensional state input in GAIL framework.
method Global encoder and reward penalization mechanism.
result Significant performance improvement in low-level and high-dimensional tasks.

This work learns latent representations to speed up exploration in complex environments.

problem Challenging exploration in high-dimensional state and action spaces with sparse rewards.
method Representation learning using prior experience to learn effective latent representations.
result Learned latent representations reduce the dimensionality of the search space for effective exploration.

Unified approach combines reward maximization and empowerment for RL.

problem Combining reward maximization and empowerment for reinforcement learning.
method Unified Bellman optimality principle for empowered reward maximization.
result Unified approach leads to improved initial and competitive final performance.

Abstract MDPs enable strategic exploration and fast reward transfer in complex environments.

problem Challenging to learn accurate MDPs for high-dimensional states.
method Learn an abstract MDP over low-dimensional coarse states, using an abstraction function.
result Achieves superhuman performance on Pitfall! and higher reward with fewer samples.

Direct approach for handling contextual bandits with latent state dynamics.

problem Handling contextual bandits with latent state dynamics, especially when rewards depend on posterior probabilities of hidden states.
method Direct reduction to standard linear contextual bandits, extended analysis of HMM parameters, periodic update of reward-model parameters.
result Periodic update of reward-model parameters allows handling complex dependencies in hidden states.

B-REX efficiently learns Atari game policies from pixel inputs using Bayesian methods.

problem Learning reward functions from visual inputs with uncertainty and safety considerations.
method Bayesian Reward Extrapolation (B-REX) using successor features and preferences.
result B-REX generates posterior samples efficiently, enabling high-confidence performance bounds.

New model for bandit problem with linear rewards and side information.

problem Hidden Markovian bandit problem with linear rewards and side information.
method Presented a model and algorithm with regret analysis for the problem.
result Logarithmic regret achieved even in high-dimensional problems with structural side information.

Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.

problem Reward-free learning in high-dimensional, continuous-control domains.
method Maximum Entropy POLicy optimization (MEPOL) algorithm that maximizes a non-parametric state entropy estimate.
result MEPOL learns a maximum-entropy exploration policy that facilitates learning various reward-based tasks.

The paper introduces a method for multi-agent reinforcement learning to coordinate exploration.

problem Sparse rewards in multi-agent settings lead to independent exploration.
method Designing intrinsic rewards that encourage coordination and developing a hierarchical policy.
result The approach accelerates and improves exploration in cooperative multi-agent settings.

Paper addresses reward learning issues in RL, improving both under- and over-estimation.

problem Reward learning from data can lead to reward delusions or underestimation, causing unintended behaviors.
method Connects reward learning to positive-unlabeled (PU) learning and applies a large-scale PU learning algorithm.
result Improves both GAIL and supervised reward learning without additional assumptions.

PHE adds pseudo-rewards to history to minimize regret in stochastic bandits.

problem Minimizing cumulative regret in stochastic multi-armed bandits.
method PHE algorithm that adds O(t)O(t) i.i.d. pseudo-rewards to history and pulls the best arm based on the perturbed history.
result Near-optimal regret bounds derived for PHE.

Paper analyzes AIRL in high-dimensional spaces using random matrix theory.

problem AIRL's performance challenges in high-dimensional environments.
method Examined the rank of the matrix derived from transition matrix, applied random matrix theory.
result High-dimensional scenarios reveal transfer limitations not inherent to AIRL framework.

STAR framework reduces OPE variance by distilling complex problems into discrete ARPs.

problem High variance and bias in off-policy evaluation methods.
method STAR framework that includes various OPE estimators and leverages state abstraction.
result Predictions from ARPs estimated from off-policy data are asymptotically correct.

We formalize and decompose reinforcement learning problems with exogenous state variables and rewards.

problem Exogenous state variables and rewards slow down reinforcement learning.
method Formalized exogenous state variables and rewards, decomposed MDPs, derived variance-covariance condition, developed algorithms.
result Monte Carlo policy evaluation on the endogenous MDP is accelerated compared to using the full MDP.

Algorithm optimizes decision-making in unknown MDPs with minimal regret.

problem Optimizing decision-making in unknown discrete MDPs with bounded expected shortest path.
method Developed BUCRL{} algorithm achieving ildeO(DSAT) ilde{\mathcal{O}}(\sqrt{DSAT}) regret.
result First polynomial time Bayesian algorithm for unknown MDPs with high probability worst-case regret.

Successor Options discovers reusable skills using landmark states.

problem Discovering reusable skills in reinforcement learning.
method Leverages Successor Representations to build a state space model and learns intra-option policies using a novel pseudo-reward.
result Demonstrates the approach's efficacy on grid-worlds and high-dimensional robotic control environments.

The paper tackles reinforcement learning with exogenous variables and rewards.

problem Exogenous state variables and rewards slow reinforcement learning by introducing uncontrolled variation.
method Formalizes exogenous state variables and rewards, decomposes MDP into exogenous and endogenous components, and introduces algorithms to discover these components.
result Optimal policies for the endogenous MDP are also optimal for the original MDP, but the endogenous MDP is easier to solve due to reduced variance.

SQIL uses a simple reward strategy to encourage long-horizon imitation of expert demonstrations.

problem Challenges in imitation learning with high-dimensional, continuous observations and unknown dynamics.
method Imitates expert demonstrations by providing a constant reward of +1 for matching actions in demonstrated states, and 0 for all others.
result Empirically outperforms behavioral cloning and achieves competitive results compared to GAIL.

DeepMDP simplifies complex observations into continuous latent states.

problem Learning from high-dimensional observations in reinforcement learning.
method Trains a DeepMDP model that predicts rewards and next latent states.
result Optimization of DeepMDP objectives ensures quality of latent space and environment model.

New method learns Atari game Montezuma's Revenge from a single demonstration.

problem Learning from sparse rewards in complex exploration tasks.
method Maximizing rewards directly from a single demonstration state, combined with off-the-shelf reinforcement learning.
result Trained agent achieves high-score of 74,500 in Montezuma's Revenge.

Method learns near-optimal rewards and policies from expert examples.

problem Learning reward and policy from expert demonstrations under unknown dynamics.
method Generative adversarial networks with empowerment-regularized maximum-entropy inverse reinforcement learning.
result Method learns near-optimal rewards and policies that generalize well.

DeepSynth synthesizes automata to guide deep RL agents through sparse, non-Markovian rewards.

problem Training deep RL agents with sparse, non-Markovian rewards and unknown high-level objectives.
method Employing a novel algorithm for synthesizing compact automata to uncover sequential structure from trace data.
result Reduces the number of iterations required for policy synthesis by two orders of magnitude and improves scalability.

New method learns high-quality Laplacian representations for reinforcement learning.

problem Lack of accurate Laplacian representations in large or continuous state spaces.
method Reformulated spectral graph drawing objective to have eigenvectors as unique global minimizer.
result Learned Laplacian representations more faithfully approximate the ground truth.

New method converts natural language commands into reward functions for robots.

problem Creating effective reward functions for autonomous machines.
method Language-conditioned reward learning (LC-RL) using inverse reinforcement learning.
result Model learns transferable rewards from natural language commands.

Develops HMRL for sparse reward RL problems, improving meta policy efficiency and transferability.

problem Difficulty in learning meta policies for sparse reward RL problems.
method Hyper-Meta RL framework with cross-environment meta state embedding and shaped meta reward.
result Improves meta policy generalization and efficiency for sparse reward RL problems.

VICE uses events to define rewards without expert demonstrations.

problem Designing effective reward functions for reinforcement learning.
method VICE generalizes inverse reinforcement learning to use event probabilities.
result VICE achieves high performance on complex tasks with high-dimensional observations.

In this paper we consider the problem of online stochastic optimization of a locally smooth function under bandit feedback. We introduce the high-confidence tree (HCT) algorithm, a novel any-time X\mathcal{X}-armed bandit algorithm, and derive regret bounds matching the performance of existing state-of-the-art in term…

2014-02-04abs ↗pdf ↗

Extends Hindsight Experience Replay to learn from multiple reward functions.

problem Learning policies for multiple reward functions without trial runs.
method Develops a method to use a reward function to calculate rewards for states not encountered, and learns policies for multiple reward functions.
result A single policy can generalize across all linear combinations of multi-objective rewards.