Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

3571106141 · Jun 202019922001200920182026
48 results for reward definition

VICE uses events to define rewards without expert demonstrations.

problem Designing effective reward functions for reinforcement learning.
method VICE generalizes inverse reinforcement learning to use event probabilities.
result VICE achieves high performance on complex tasks with high-dimensional observations.

New definition resolves ambiguity in non-stationary bandit classification.

problem Ambiguity in classifying non-stationary bandits using existing definitions.
method Introducing a formal definition that resolves ambiguity and provides a unified approach.
result Unified approach applicable to both Bayesian and frequentist formulations, resolves classification issues.

The paper introduces a health-informed policy gradient method for multi-agent reinforcement learning.

problem Optimizing joint reward functions in multi-agent systems with varying agent health.
method Health-informed credit assignment in a multi-agent proximal policy optimization algorithm.
result Significant improvement in learning performance compared to traditional methods.

VASE uses Bayesian neural networks to improve exploration in sparse reward environments.

problem Exploration in environments with continuous control and sparse rewards.
method VASE uses a Bayesian neural network model of the environment dynamics and variational inference to alternately update the model's accuracy and policy.
result VASE outperforms other surprise-based exploration techniques in continuous control sparse reward environments.

A new framework assesses financial and ESG risks for sustainable investing.

problem Measuring risk and reward in sustainable investing considering environmental, social, and governance factors.
method Proposes axiomatic definitions for ESG-coherent risk measures and reward-risk ratios based on bivariate random variables.
result Empirical analysis ranks stocks using the proposed measures.

The paper studies reward concentration in MDPs, covering asymptotic and non-asymptotic settings.

problem Reward concentration in Markov Decision Processes (MDPs).
method Unified approach to reward concentration in MDPs, including asymptotic and non-asymptotic bounds.
result Rate-equivalent definitions of regret for learning policies.

New RL approach uses future state and action visitation measures for better exploration.

problem Improving exploration in reinforcement learning.
method Intrinsic reward based on future state and action visitation measures, using contraction operators.
result Policies achieve good state-action space coverage and high performance.

Paper proposes a method to learn and exceed expert demonstrations in unknown reward environments.

problem Learning to outperform expert demonstrations in unknown reward environments.
method A novel concurrent reward and action policy learning approach with a stereo utility definition.
result The proposed method can outperform expert demonstrations in various environments.

The paper introduces a new intrinsic reward method for exploration in reinforcement learning.

problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.

New algorithm tackles CMAB with filtered feedback, achieving O(ln(n))\mathcal{O}(\ln(n)) regret.

problem Sequential search and detection problems with hidden true rewards.
method Robust-F-CUCB algorithm, balancing exploration and exploitation.
result Upper confidence bound algorithm with O(ln(n))\mathcal{O}(\ln(n)) regret bound.

Combines experience replay and exploration for better agent performance.

problem Improving exploration efficiency and robustness in reinforcement learning.
method Integrates Intrinsic Rewards with Prioritized Oversampled Experience Replay (POER).
result Achieves better agent performance and sample efficiency compared to PPO/RND.

Investigates risk measures for DC pension decumulation.

problem Develop optimal decumulation strategies for DC plan holders.
method Formulates decumulation as a control problem, studies risk measures (expected shortfall, linear shortfall, probability of shortfall).
result Optimal controls for expected reward and expected shortfall are identical to those for expected reward and linear shortfall.

Capital allocation principles are used in various contexts in which a risk capital or a cost of an aggregate position has to be allocated among its constituent parts. We study capital allocation principles in a performance measurement framework. We introduce the notation of suitability of allocations for performance me…

2013-01-23abs ↗pdf ↗

Paper estimates risks in MDPs using state lumping and SAT, showing its effectiveness.

problem Estimating risks in Markov decision processes with state augmentation.
method State augmentation transformation, isotopic states, and state lumping.
result SAT and state lumping effectively estimate mean-variance and exponential utility risks.

FRONT optimizes decisions with interference, reducing regret over time.

problem Short-sighted policies in online decision-making due to ignoring interference.
method FRONT considers long-term impacts of decisions, using exploratory and exploitative strategies.
result FRONT achieves sublinear regret in both immediate and consequential impacts.

Improved regret bound for multinomial logistic bandits with non-linearity.

problem Maximizing rewards in multinomial logistic bandits with non-linear feedback.
method Extended the definition of κκ_* to multinomial setting and proposed an efficient algorithm.
result Minimax-optimal regret bound of O~(RdKT/κ) \smash{\widetilde{\mathcal{O}}( R d \sqrt{ {KT}/{κ_*}} ) } , improving over existing guarantees.

The paper develops a learning framework for diverse legged robots.

problem General and autonomous learning of core skills in locomotion.
method Data-efficient, off-policy multi-task RL algorithm with semantically identical reward functions.
result The same algorithm can learn diverse and reusable locomotion skills across different legged robots.

Improved Thompson Sampling using fractional posteriors achieves better regret bounds.

problem Optimizing regret in stochastic multi-armed bandit problems.
method Using α\alpha-posterior distributions, derived frequentist regret bounds.
result Instance-dependent and instance-independent regret bounds established.

Paper tackles outlier detection in multi-armed bandits, achieving high accuracy with reduced exploration costs.

problem Detecting outlier arms in multi-armed bandit settings.
method Proposes GOLD algorithm based on upper confidence bounds to identify generic outlier arms.
result Achieves 98% accuracy with 83% reduction in exploration cost compared to state-of-the-art techniques.

Study quantile multi-armed bandits for identifying the best arm with a specified quantile level.

problem Identifying the arm with the highest quantile in multi-armed bandits with private rewards.
method Proposed a (non-private) and differentially private successive elimination algorithms for best-arm identification.
result The proposed algorithms are essentially optimal for quantile bandit problems, with finite sample complexity even for distributions with infinite support-size.

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

Paper addresses reward learning issues in RL, improving both under- and over-estimation.

problem Reward learning from data can lead to reward delusions or underestimation, causing unintended behaviors.
method Connects reward learning to positive-unlabeled (PU) learning and applies a large-scale PU learning algorithm.
result Improves both GAIL and supervised reward learning without additional assumptions.

The study categorizes reward errors in reinforcement learning, finding some can be beneficial.

problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.

Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.

problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

Proposes a method to boost deep reinforcement learning with sparse rewards.

problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.

Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.

problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.