New RL formulation for maximizing maximum reward in molecule generation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Scalable model for slate recommendation learns reward probabilities.
New RL approach uses future state and action visitation measures for better exploration.
We propose a new complexity measure for Markov decision processes (MDPs), the maximum expected hitting cost (MEHC). This measure tightens the closely related notion of diameter [JOA10] by accounting for the reward structure. We show that this parameter replaces diameter in the upper bound on the optimal value span of a…
Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.
New method learns multiple reward functions for complex tasks.
We introduce a new online learning framework where, at each trial, the learner is required to select a subset of actions from a given known action set. Each action is associated with an energy value, a reward and a cost. The sum of the energies of the actions selected cannot exceed a given energy budget. The goal is to…
New algorithms improve linear bandit performance with low computation.
NeuralRBMLE tackles explore-exploit trade-offs in contextual bandits with neural networks.
A new method recovers rewards from behavior policies using classification and regression.
This work tackles the challenge of aligning generative models without explicit reward signals.
Novel algorithm reduces computational burden in IRL with finite-time guarantees.
While most approaches to the problem of Inverse Reinforcement Learning (IRL) focus on estimating a reward function that best explains an expert agent's policy or demonstrated behavior on a control task, it is often the case that such behavior is more succinctly represented by a simple reward combined with a set of hard…
Improved GP bandit algorithms for noiseless, varying noise, and RKHS norms.
We consider the problem of learning from demonstrated trajectories with inverse reinforcement learning (IRL). Motivated by a limitation of the classical maximum entropy model in capturing the structure of the network of states, we propose an IRL model based on a generalized version of the causal entropy maximization pr…
No real-world reward function is perfect. Sensory errors and software bugs may result in RL agents observing higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formal…
Designers of AI agents often iterate on the reward function in a trial-and-error process until they get the desired behavior, but this only guarantees good behavior in the training environment. We propose structuring this process as a series of queries asking the user to compare between different reward functions. Thus…
Unified meta algorithms estimate various distribution functionals in infinite-armed bandits.
Diversified risk parity strategies outperform equally-weighted portfolios in various asset universes.
Paper generalizes reward distribution in multi-armed bandits with temporally-partitioned rewards.
Reward augmented maximum likelihood (RAML), a simple and effective learning framework to directly optimize towards the reward function in structured prediction tasks, has led to a number of impressive empirical successes. RAML incorporates task-specific reward by performing maximum-likelihood updates on candidate outpu…
Paper tackles offline preference-based RL with human feedback.
New algorithm reduces regret in infinitely many-armed bandits with decreasing rewards.
A new algorithm balances global reward and group constraints in federated multi-armed bandits.
We implement momentum strategies using reward-risk measures as ranking criteria based on classical tempered stable distribution. Performances and risk characteristics for the alternative portfolios are obtained in various asset classes and markets. The reward-risk momentum strategies with lower volatility levels outper…
Text generation is a crucial task in NLP. Recently, several adversarial generative models have been proposed to improve the exposure bias problem in text generation. Though these models gain great success, they still suffer from the problems of reward sparsity and mode collapse. In order to address these two problems, …
New approach for reward-free exploration reduces estimation error.
Inspired by the Reward-Biased Maximum Likelihood Estimate method of adaptive control, we propose RBMLE -- a novel family of learning algorithms for stochastic multi-armed bandits (SMABs). For a broad range of SMABs including both the parametric Exponential Family as well as the non-parametric sub-Gaussian/Exponential f…
Unified LP framework for offline reward learning from human demonstrations and feedback.
We study the Combinatorial Pure Exploration problem with Continuous and Separable reward functions (CPE-CS) in the stochastic multi-armed bandit setting. In a CPE-CS instance, we are given several stochastic arms with unknown distributions, as well as a collection of possible decisions. Each decision has a reward accor…
Reinforcement learning agents are prone to undesired behaviors due to reward mis-specification. Finding a set of reward functions to properly guide agent behaviors is particularly challenging in multi-agent scenarios. Inverse reinforcement learning provides a framework to automatically acquire suitable reward functions…
We consider a problem of learning the reward and policy from expert examples under unknown dynamics. Our proposed method builds on the framework of generative adversarial networks and introduces the empowerment-regularized maximum-entropy inverse reinforcement learning to learn near-optimal rewards and policies. Empowe…
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
The paper shows how policy regularization acts like an adversary to improve robustness.
Unified framework for estimating reward functions in competitive games.
In many settings (e.g., robotics) demonstrations provide a natural way to specify tasks; however, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for the tasks, such as rewards or policies, can be safely composed and/or do not explicitly capture history dependen…
The paper develops a reinforcement learning model to estimate ad impact considering delayed and cumulative effects.
Improved exploration methods for reinforcement learning with reduced sample complexity.
New algorithms improve contextual bandits with neural networks and energy models.
A new framework isolates exploration challenges in RL without explicit rewards.
Actor critic methods with sparse rewards in model-based deep reinforcement learning typically require a deterministic binary reward function that reflects only two possible outcomes: if, for each step, the goal has been achieved or not. Our hypothesis is that we can influence an agent to learn faster by applying an ext…
We consider stochastic multi-armed bandit problems with complex actions over a set of basic arms, where the decision maker plays a complex action rather than a basic arm in each round. The reward of the complex action is some function of the basic arms' rewards, and the feedback observed may not necessarily be the rewa…
The paper analyzes the reward improvement of aligned policies in large language models.
Develops statistical framework for resolving reward function ambiguity in inverse reinforcement learning.
In reinforcement learning the Q-values summarize the expected future rewards that the agent will attain. However, they cannot capture the epistemic uncertainty about those rewards. In this work we derive a new Bellman operator with associated fixed point we call the `knowledge values'. These K-values compress both the …
Enhances portfolio performance using deep reinforcement learning and future rewards.
Study shows fast rates for inverse reinforcement learning with linear rewards.