Paper proposes a method to learn goal-reaching behaviors from scratch using imitation learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Sparse reward problems are one of the biggest challenges in Reinforcement Learning. Goal-directed tasks are one such sparse reward problems where a reward signal is received only when the goal is reached. One promising way to train an agent to perform goal-directed tasks is to use Hindsight Learning approaches. In thes…
Eikonal-Constrained QRL improves goal-reaching in reinforcement learning.
We determine the optimal strategies for purchasing term life insurance and for investing in a risky financial market in order to maximize the probability of reaching a bequest goal while consuming from an investment account. We extend Bayraktar and Young (2015) by allowing the individual to purchase term life insurance…
We determine the optimal strategy for investing in a Black-Scholes market in order to maximize the probability that wealth at death meets a bequest goal , a type of goal-seeking problem, as pioneered by Dubins and Savage (1965, 1976). The individual consumes at a constant rate , so the level of wealth required fo…
Autonomous agents that must exhibit flexible and broad capabilities will need to be equipped with large repertoires of skills. Defining each skill with a manually-designed reward function limits this repertoire and imposes a manual engineering burden. Self-supervised agents that set their own goals can automate this pr…
C-Learning estimates reachability over time to solve multi-goal tasks.
This work speeds up imitation learning for reaching any goal.
The paper examines how background risk affects portfolio selection and optimal reinsurance design.
Method maps state space using landmarks for universal goal reaching.
New approach to goal-based investing using hedging and reinforcement learning.
LEXA learns to discover and achieve goals in unseen environments.
We consider the problem of how an individual can use term life insurance to maximize the probability of reaching a given bequest goal, an important problem in financial planning. We assume that the individual buys instantaneous term life insurance with a premium payable continuously. By contrast with Bayraktar et al. (…
In this paper, we consider three problems related to survival, growth, and goal reaching maximization of an investment portfolio with proportional net cash flow. We solve the problems in a market constrained due to borrowing prohibition. To solve the problems, we first construct an auxiliary market and then apply the d…
HGG generates goals to improve sample efficiency in robotic tasks.
This work improves imitation learning and goal-conditioned RL by estimating value densities.
We determine how an individual can use life insurance to meet a bequest goal. We assume that the individual's consumption is met by an income, such as a pension, life annuity, or Social Security. Then, we consider the wealth that the individual wants to devote towards heirs (separate from any wealth related to the afor…
A method for setting up an automatic curriculum for reinforcement learning tasks.
Physics-informed GCRL tackles sparse feedback learning with hybrid dynamics.
A new method for robot manipulation tasks using imagined object goals.
While reinforcement learning (RL) has the potential to enable robots to autonomously acquire a wide range of skills, in practice, RL usually requires manual, per-task engineering of reward functions, especially in real world settings where aspects of the environment needed to compute progress are not directly accessibl…
This work tackles long-term visual planning by goal-conditioned hierarchical predictors.
Framework discovers sub-goals for better exploration in RL tasks.
Single neural network predicts ImageNet model parameters for faster training.
Deep reinforcement learning has recently gained a focus on problems where policy or value functions are independent of goals. Evidence exists that the sampling of goals has a strong effect on the learning performance, but there is a lack of general mechanisms that focus on optimizing the goal sampling process. In this …
Hierarchical Foresight improves robot vision tasks by planning long-term goals.
Learning to control an environment without hand-crafted rewards or expert data remains challenging and is at the frontier of reinforcement learning research. We present an unsupervised learning algorithm to train agents to achieve perceptually-specified goals using only a stream of observations and actions. Our agent s…
For an autonomous agent to fulfill a wide range of user-specified goals at test time, it must be able to learn broadly applicable and general-purpose skill repertoires. Furthermore, to provide the requisite level of generality, these skills must handle raw sensory input such as images. In this paper, we propose an algo…
Automatically learns dynamical distances for efficient reinforcement learning.
Consider mutli-goal tasks that involve static environments and dynamic goals. Examples of such tasks, such as goal-directed navigation and pick-and-place in robotics, abound. Two types of Reinforcement Learning (RL) algorithms are used for such tasks: model-free or model-based. Each of these approaches has limitations.…
MpFL models clients as strategic players to reach equilibrium with less communication.
A key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization. To this end, we introduce universal planning networks (UPN). UPNs embed differentiable planning within a goal-directed policy. This planning computation unrolls a for…
Being able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However …
Goals for reinforcement learning problems are typically defined through hand-specified rewards. To design such problems, developers of learning algorithms must inherently be aware of what the task goals are, yet we often require agents to discover them on their own without any supervision beyond these sparse rewards. W…
GOIMDA selects inputs to maximize expected influence on a goal functional, reducing data acquisition needs.
Directed exploration improves reinforcement learning efficiency and robustness.
Our goal is to show the beauty and power of Alexandrov geometry by reaching interesting applications and theorems with a minimum of preparation. The topics include 1. Reshetnyak's gluing theorem, 2. Estimates on the number of collisions in billiards, 3. Reshetnyak's majorization theorem, 4. Hadamard--Cartan globalizati…
Exploration is a difficult challenge in reinforcement learning and even recent state-of-the art curiosity-based methods rely on the simple epsilon-greedy strategy to generate novelty. We argue that pure random walks do not succeed to properly expand the exploration area in most environments and propose to replace singl…
A method uses RL to learn abstractions for planning, improving robot navigation and manipulation tasks.
The paper develops methods to create fair and transferable representations without subgroup discrimination.
Branch-and-bound (BnB) algorithms are widely used to solve combinatorial problems, and the performance crucially depends on its branching heuristic.In this work, we consider a typical problem of maximum common subgraph (MCS), and propose a branching heuristic inspired from reinforcement learning with a goal of reaching…
New RL framework learns task completion without prior knowledge.
Explains how machine learning models can be biased and presents interactive plots to visualize bias.
The globalization feeded by the technology explosion that begans in the end of the last century, started the world to change faster every day. The only today's certain is the tomorrow's uncertain. Risk is defined as uncertain where one or many causes composed of ocurrence probality can generate an impact or consequence…
Inferring a person's goal from their behavior is an important problem in applications of AI (e.g. automated assistants, recommender systems). The workhorse model for this task is the rational actor model - this amounts to assuming that people have stable reward functions, discount the future exponentially, and construc…
For a Lie group G and a smooth manifold W, we study the difference between smooth actions of G on W and bundles over the classifying space of G with fiber W and structure group Diff(W). In particular, we exhibit smooth manifold bundles over BSU(2) that are not induced by an action. The main tool for reaching this goal …
The paper optimizes insurance purchases for financial goals.
Study completes braid index determination for all pretzel links.