VICE uses events to define rewards without expert demonstrations.
problem Designing effective reward functions for reinforcement learning.
method VICE generalizes inverse reinforcement learning to use event probabilities.
result VICE achieves high performance on complex tasks with high-dimensional observations.
New definition resolves ambiguity in non-stationary bandit classification.
problem Ambiguity in classifying non-stationary bandits using existing definitions.
method Introducing a formal definition that resolves ambiguity and provides a unified approach.
result Unified approach applicable to both Bayesian and frequentist formulations, resolves classification issues.
The paper introduces a health-informed policy gradient method for multi-agent reinforcement learning.
problem Optimizing joint reward functions in multi-agent systems with varying agent health.
method Health-informed credit assignment in a multi-agent proximal policy optimization algorithm.
result Significant improvement in learning performance compared to traditional methods.
New method uses simple sensor intentions to learn complex tasks.
problem Defining reward schemes for exploration in robotic systems.
method Introduce simple sensor intentions (SSIs) to define auxiliary tasks.
result Learning system can solve complex robotic tasks using only raw sensor streams.
VASE uses Bayesian neural networks to improve exploration in sparse reward environments.
problem Exploration in environments with continuous control and sparse rewards.
method VASE uses a Bayesian neural network model of the environment dynamics and variational inference to alternately update the model's accuracy and policy.
result VASE outperforms other surprise-based exploration techniques in continuous control sparse reward environments.
A new framework assesses financial and ESG risks for sustainable investing.
problem Measuring risk and reward in sustainable investing considering environmental, social, and governance factors.
method Proposes axiomatic definitions for ESG-coherent risk measures and reward-risk ratios based on bivariate random variables.
result Empirical analysis ranks stocks using the proposed measures.
New concept: reward hacking, where optimizing a flawed reward function can hurt performance.
problem Optimizing imperfect reward functions leads to poor performance.
method Formal definition and analysis of reward hacking, examining conditions for unhackability.
result Reward functions are usually hackable, making it hard to align AI with human values.
The paper studies reward concentration in MDPs, covering asymptotic and non-asymptotic settings.
problem Reward concentration in Markov Decision Processes (MDPs).
method Unified approach to reward concentration in MDPs, including asymptotic and non-asymptotic bounds.
result Rate-equivalent definitions of regret for learning policies.
New RL approach uses future state and action visitation measures for better exploration.
problem Improving exploration in reinforcement learning.
method Intrinsic reward based on future state and action visitation measures, using contraction operators.
result Policies achieve good state-action space coverage and high performance.
Unified definition of hallucinations in language models.
problem Persistent hallucinations despite mitigation efforts.
method Unified definition of hallucination as inaccurate world modeling.
result Unified framework distinguishes hallucinations from other errors.
Paper proposes a method to learn and exceed expert demonstrations in unknown reward environments.
problem Learning to outperform expert demonstrations in unknown reward environments.
method A novel concurrent reward and action policy learning approach with a stereo utility definition.
result The proposed method can outperform expert demonstrations in various environments.
Unified DP framework for multi-armed bandits with privacy cost.
problem Privacy in multi-armed bandit algorithms.
method Differential privacy framework, unified graphical model, lower bounds derivation.
result Regret is increased by a factor dependent on privacy level ε.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.
New method tackles bandit problem with varying rewards over time.
problem Varying rewards in bandit problems over time.
method Gaussian processes for estimation and planning.
result Improved computational efficiency through optimistic planning.
New algorithm tackles CMAB with filtered feedback, achieving O ( ln ( n ) ) \mathcal{O}(\ln(n)) O ( ln ( n )) regret.
problem Sequential search and detection problems with hidden true rewards.
method Robust-F-CUCB algorithm, balancing exploration and exploitation.
result Upper confidence bound algorithm with O ( ln ( n ) ) \mathcal{O}(\ln(n)) O ( ln ( n )) regret bound. Combines experience replay and exploration for better agent performance.
problem Improving exploration efficiency and robustness in reinforcement learning.
method Integrates Intrinsic Rewards with Prioritized Oversampled Experience Replay (POER).
result Achieves better agent performance and sample efficiency compared to PPO/RND.
Investigates risk measures for DC pension decumulation.
problem Develop optimal decumulation strategies for DC plan holders.
method Formulates decumulation as a control problem, studies risk measures (expected shortfall, linear shortfall, probability of shortfall).
result Optimal controls for expected reward and expected shortfall are identical to those for expected reward and linear shortfall.
Unified theory for UCB policies in total and max bandit problems.
problem Order optimality of UCB policies in max bandit problems.
method Unified definition of UCB policy using oracle quantity and failure count.
result UCB policies are order optimal in both total and max bandit problems.
Vroom optimizes in unpredictable conditions without derivatives.
problem Optimizing in non-stationary, adversarial environments.
method Zeroth-order online learning with vanishing regret.
result Achieves favorable rates in stochastic settings.
Capital allocation principles are used in various contexts in which a risk capital or a cost of an aggregate position has to be allocated among its constituent parts. We study capital allocation principles in a performance measurement framework. We introduce the notation of suitability of allocations for performance me…
Paper estimates risks in MDPs using state lumping and SAT, showing its effectiveness.
problem Estimating risks in Markov decision processes with state augmentation.
method State augmentation transformation, isotopic states, and state lumping.
result SAT and state lumping effectively estimate mean-variance and exponential utility risks.
Contextual bandits study how reward variance affects regret bounds.
problem Investigating how small reward variance impacts regret bounds in contextual bandits.
method Analyzing two types of adversaries and function approximation complexities.
result Regret bounds are influenced by the eluder dimension and reward variance.
ERTS uses Thompson sampling for Gaussian entropic risk bandits, achieving regret bounds.
problem Risk in decision making complicates reward maximization in MAB problems.
method ERTS (Entropic Risk Thompson Sampling) using Thompson sampling with an entropic risk measure.
result Regret bounds for ERTS under entropic risk measure provided.
We introduce various quantitative and mathematical definitions for price momentum of financial instruments. The price momentum is quantified with velocity and mass concepts originated from the momentum in physics. By using the physical momentum of price as a selection criterion, the weekly contrarian strategies are imp…
FRONT optimizes decisions with interference, reducing regret over time.
problem Short-sighted policies in online decision-making due to ignoring interference.
method FRONT considers long-term impacts of decisions, using exploratory and exploitative strategies.
result FRONT achieves sublinear regret in both immediate and consequential impacts.
Improved regret bound for multinomial logistic bandits with non-linearity.
problem Maximizing rewards in multinomial logistic bandits with non-linear feedback.
method Extended the definition of κ ∗ κ_* κ ∗ to multinomial setting and proposed an efficient algorithm. result Minimax-optimal regret bound of O ~ ( R d K T / κ ∗ ) \smash{\widetilde{\mathcal{O}}( R d \sqrt{ {KT}/{κ_*}} ) } O ( R d K T / κ ∗ ) , improving over existing guarantees. HGG generates goals to improve sample efficiency in robotic tasks.
problem Efficiency in reinforcement learning with sparse reward signals.
method Generates valuable hindsight goals for reinforcement learning.
result Significantly improved sample efficiency over HER.
This paper explores RL for financial trading systems.
problem Developing automatic Financial Trading Systems (FTFs).
method Explains RL concepts and applies them to FTFs.
result Illustrates RL algorithms for FTFs.
The paper develops a learning framework for diverse legged robots.
problem General and autonomous learning of core skills in locomotion.
method Data-efficient, off-policy multi-task RL algorithm with semantically identical reward functions.
result The same algorithm can learn diverse and reusable locomotion skills across different legged robots.
A new algorithm improves regret bounds for contextual bandits.
problem Complexity of missing data in multi-armed bandits.
method Doubly Robust (DR) Thompson Sampling with contexts.
result Improved regret bound with i l d e O ( d T ) ilde{O}(d\sqrt{T}) i l d e O ( d T ) . Improved Thompson Sampling using fractional posteriors achieves better regret bounds.
problem Optimizing regret in stochastic multi-armed bandit problems.
method Using α \alpha α -posterior distributions, derived frequentist regret bounds. result Instance-dependent and instance-independent regret bounds established.
Paper tackles outlier detection in multi-armed bandits, achieving high accuracy with reduced exploration costs.
problem Detecting outlier arms in multi-armed bandit settings.
method Proposes GOLD algorithm based on upper confidence bounds to identify generic outlier arms.
result Achieves 98% accuracy with 83% reduction in exploration cost compared to state-of-the-art techniques.
Study differential privacy in contextual linear bandits.
problem Maximizing rewards in a contextual bandit problem with private data.
method Adopted joint differential privacy, converted classic linear-UCB to joint-differentially private algorithm.
result First lower bound on additional regret for private algorithms.
Study quantile multi-armed bandits for identifying the best arm with a specified quantile level.
problem Identifying the arm with the highest quantile in multi-armed bandits with private rewards.
method Proposed a (non-private) and differentially private successive elimination algorithms for best-arm identification.
result The proposed algorithms are essentially optimal for quantile bandit problems, with finite sample complexity even for distributions with infinite support-size.
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.
The paper defines price sensitivity and liquidity in CFMMs and links it to curvature.
problem Understanding the relationship between CFMM curvature and market performance.
method Proposes a definition of price sensitivity and liquidity, and links it to CFMM curvature.
result Curvature of CFMMs affects market performance and liquidity provider incentives.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.
Paper addresses reward learning issues in RL, improving both under- and over-estimation.
problem Reward learning from data can lead to reward delusions or underestimation, causing unintended behaviors.
method Connects reward learning to positive-unlabeled (PU) learning and applies a large-scale PU learning algorithm.
result Improves both GAIL and supervised reward learning without additional assumptions.
This work analyzes the value of future reward information in RL.
problem Analyzing the impact of knowing future rewards in reinforcement learning.
method Competitive analysis and worst-case reward distribution.
result Exact ratios between standard RL agents and those with future-reward lookahead.
Method learns reward functions that are independently obtainable and sum to original reward.
problem Learning reward functions that are independent and meaningful.
method Defining independent obtainability and optimizing a novel objective function.
result Learned reward functions generalize well to modified environments and have optimal policies.
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.
Self-supervised reward prediction improves RL in sparse reward settings.
problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.
Reward models need more than just accuracy for effective RLHF.
problem The effectiveness of reward models in RLHF is not fully understood.
method An optimization perspective to evaluate reward models.
result Reward models with low reward variance can lead to a flat optimization landscape, hindering performance.
Proposes a method to boost deep reinforcement learning with sparse rewards.
problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.
Paper proposes CVaR-TS for risk-constrained MAB problems.
problem Risk in decision-making complicates reward maximization in MAB problems.
method Risk measure CVaR is used, and Thompson Sampling is adapted for CVaR.
result CVaR-TS outperforms other L/UCB-based algorithms in risk-constrained MAB settings.
Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.
problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.
Reward estimation improves model-free RL performance under corrupted rewards.
problem Handling corrupted or stochastic rewards in reinforcement learning.
method Using an estimator for both rewards and value functions.
result Improves performance in various noise types and environments.