PCHID improves sample efficiency in reinforcement learning tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Rewriting history improves RL algorithms for solving multiple tasks.
Generalized Hindsight improves RL by reusing data from one task for another.
Dynamical-VAE learns causal dynamics from POMDPs using future information.
Dealing with sparse rewards is a longstanding challenge in reinforcement learning. The recent use of hindsight methods have achieved success on a variety of sparse-reward tasks, but they fail on complex tasks such as stacking multiple blocks with a robot arm in simulation. Curiosity-driven exploration using the predict…
New method uses hindsight to make exploration robust in stochastic environments.
New algorithms learn POMDPs efficiently with hindsight observability.
Reinforcement Learning(RL) with sparse rewards is a major challenge. We propose \emph{Hindsight Trust Region Policy Optimization}(HTRPO), a new RL algorithm that extends the highly successful TRPO algorithm with \emph{hindsight} to tackle the challenge of sparse rewards. Hindsight refers to the algorithm's ability to l…
Sparse reward problems are one of the biggest challenges in Reinforcement Learning. Goal-directed tasks are one such sparse reward problems where a reward signal is received only when the goal is reached. One promising way to train an agent to perform goal-directed tasks is to use Hindsight Learning approaches. In thes…
We study T. Cover's rebalancing option (Ordentlich and Cover 1998) under discrete hindsight optimization in continuous time. The payoff in question is equal to the final wealth that would have accrued to a $\$1$ deposit into the best of some finite set of (perhaps levered) rebalancing rules determined in hindsight. A r…
Framework uses hindsight regret to audit marketing budget allocations.
New algorithms assign credit to past decisions based on hindsight.
Paper introduces a new regret measure for online convex optimization with smooth losses.
Efficient RL in partially observable risk-sensitive environments with hindsight observations.
Efficient algorithm for unknown linear systems with convex costs.
Interactive learning with hindsight instruction feedback achieves better performance than traditional methods.
HO2 learns options from data efficiently, improving robot manipulation tasks.
We study the control of a linear dynamical system with adversarial disturbances (as opposed to statistical noise). The objective we consider is one of regret: we desire an online control procedure that can do nearly as well as that of a procedure that has full knowledge of the disturbances in hindsight. Our main result…
The paper explores how planning with models improves credit assignment in reinforcement learning.
This paper derives a robust on-line equity trading algorithm that achieves the greatest possible percentage of the final wealth of the best pairs rebalancing rule in hindsight. A pairs rebalancing rule chooses some pair of stocks in the market and then perpetually executes rebalancing trades so as to maintain a target …
Efficient algorithm controls unknown systems with adversarial perturbations.
Goal-oriented reinforcement learning has recently been a practical framework for robotic manipulation tasks, in which an agent is required to reach a certain goal defined by a function on the state space. However, the sparsity of such reward definition makes traditional reinforcement learning algorithms very inefficien…
In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…
HL algorithms improve resource allocation in cloud environments.
New approach generates optimal disturbances for controller verification.
A new EM framework for goal-conditioned RL improves performance on sparse reward tasks.
Algorithm selects best model based on state, reducing costs.
New method improves meta-reinforcement learning efficiency.
In Hindsight Experience Replay (HER), a reinforcement learning agent is trained by treating whatever it has achieved as virtual goals. However, in previous work, the experience was replayed at random, without considering which episode might be the most valuable for learning. In this paper, we develop an energy-based fr…
Recent literature on online learning has focused on developing adaptive algorithms that take advantage of a regularity of the sequence of observations, yet retain worst-case performance guarantees. A complementary direction is to develop prediction methods that perform well against complex benchmarks. In this paper, we…
We consider dynamic pricing with many products under an evolving but low-dimensional demand model. Assuming the temporal variation in cross-elasticities exhibits low-rank structure based on fixed (latent) features of the products, we show that the revenue maximization problem reduces to an online bandit convex optimiza…
Experience replay is an important technique for addressing sample-inefficiency in deep reinforcement learning (RL), but faces difficulty in learning from binary and sparse rewards due to disproportionately few successful experiences in the replay buffer. Hindsight experience replay (HER) was recently proposed to tackle…
New method for PKM inverse dynamics second derivatives efficiently.
Reinforcement Learning (RL) algorithms can suffer from poor sample efficiency when rewards are delayed and sparse. We introduce a solution that enables agents to learn temporally extended actions at multiple levels of abstraction in a sample efficient and automated fashion. Our approach combines universal value functio…
Compared to reinforcement learning, imitation learning (IL) is a powerful paradigm for training agents to learn control policies efficiently from expert demonstrations. However, in most cases, obtaining demonstration data is costly and laborious, which poses a significant challenge in some scenarios. A promising altern…
Optimal control in changing systems without strong convexity assumptions.
This paper prices and replicates the financial derivative whose payoff at is the wealth that would have accrued to a $\$1$ deposit into the best continuously-rebalanced portfolio (or fixed-fraction betting scheme) determined in hindsight. For the single-stock Black-Scholes market, Ordentlich and Cover (1998) only p…
This work improves RL for complex robotic tasks by guiding exploration with task-specific goal distributions.
We present an adversarial active exploration for inverse dynamics model learning, a simple yet effective learning scheme that incentivizes exploration in an environment without any human intervention. Our framework consists of a deep reinforcement learning (DRL) agent and an inverse dynamics model contesting with each …
New algorithm reduces control error in systems with changing dynamics.
New method learns population dynamics from snapshots using JKO scheme and inverse optimization.
The so-called inverse problem of dynamics is about constructing a potential for a given family of curves. We observe that there is a more general way of posing the problem by making use of ideas of another inverse problem, namely the inverse problem of the calculus of variations. We critically review and clarify differ…
In E-commerce advertising, where product recommendations and product ads are presented to users simultaneously, the traditional setting is to display ads at fixed positions. However, under such a setting, the advertising system loses the flexibility to control the number and positions of ads, resulting in sub-optimal p…
This study develops a dynamic inverse optimization framework to recover hidden, time-varying preferences from observed allocation trajectories.
New algorithm combines curriculum learning with HER for complex object manipulation tasks.
New algorithm reduces performance loss in IRL with mismatched transition dynamics.
Sequential screening and dynamic regret in multi-armed bandits with arriving arms
Optimal algorithms for mixable losses in dynamic environments with reduced redundancy.