HashReward improves imitation learning in high-dimensional environments by balancing reward generation and dimensionality reduction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Models that can simulate how environments change in response to actions can be used by agents to plan and act efficiently. We improve on previous environment simulators from high-dimensional pixel observations by introducing recurrent neural networks that are able to make temporally and spatially coherent predictions f…
A new reward learning module improves imitation learning in high-dimensional environments.
New approach predicts under latent shifts using high-dimensional images.
Study on meta-reinforcement learning generalization in high-dimensional tasks.
Novel graph-spanning algorithm detects changes in high-dimensional data.
Reinforcement learning algorithms, though successful, tend to over-fit to training environments hampering their application to the real-world. This paper proposes -- a robust reinforcement learning algorithm with significant robust performance on low and high-dimensional control tasks. Ou…
Paper analyzes AIRL in high-dimensional spaces using random matrix theory.
Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces. Unlike traditional point estimate methods, randomized value functions maintain a posterior distribution over action-space values. This prevents the …
Deep RL algorithm trades high-dimensional stock portfolios.
In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent's representations during training or via use as part of an explicit planning mechanism. However, their application in practice has been limited to simplistic envi…
Study on estimating causal effects with limited data and multiple environments.
Self-supervised reward prediction improves RL in sparse reward settings.
Improved CEM for fast real-time planning in high-dimensional control tasks.
Proposes a method to improve few-shot transfer in off-dynamics RL.
Novel AMP framework for multi-environment transfer learning.
ATLAS separates invariant and transferable latent factors across diverse environments.
Study compares high-dimensional BO algorithms on 24 functions.
Teacher algorithm helps DRL learn diverse environments efficiently.
Intrinsically motivated goal exploration processes enable agents to autonomously sample goals to explore efficiently complex environments with high-dimensional continuous actions. They have been applied successfully to real world robots to discover repertoires of policies producing a wide diversity of effects. Often th…
Paper tackles efficient navigation in constrained environments using supervised and reinforcement learning.
This research tackles intervention-centric causal reasoning in learning agents by using meta-learning.
Imitation learning algorithms can be used to learn a policy from expert demonstrations without access to a reward signal. However, most existing approaches are not applicable in multi-agent settings due to the existence of multiple (Nash) equilibria and non-stationary environments. We propose a new framework for multi-…
High-dimensional always-changing environments constitute a hard challenge for current reinforcement learning techniques. Artificial agents, nowadays, are often trained off-line in very static and controlled conditions in simulation such that training observations can be thought as sampled i.i.d. from the entire observa…
Many reinforcement learning (RL) tasks provide the agent with high-dimensional observations that can be simplified into low-dimensional continuous states. To formalize this process, we introduce the concept of a DeepMDP, a parameterized latent space model that is trained via the minimization of two tractable losses: pr…
Reinforcement learning is concerned with identifying reward-maximizing behaviour policies in environments that are initially unknown. State-of-the-art reinforcement learning approaches, such as deep Q-networks, are model-free and learn to act effectively across a wide range of environments such as Atari games, but requ…
Paper explores differential privacy in high-dimensional federated learning, tackling server trustworthiness and estimation.
Paper proposes EILLS for invariant linear regression across environments.
We present a modelling framework for the investigation of prototype-based classifiers in non-stationary environments. Specifically, we study Learning Vector Quantization (LVQ) systems trained from a stream of high-dimensional, clustered data.We consider standard winner-takes-all updates known as LVQ1. Statistical prope…
Novel unsupervised MIG detectors improve signal detection in cluttered environments.
Paper develops a model-based RL framework for portfolio optimization in financial markets.
Deep RL approach improves MIS for complex environments.
Unified approach to neural network learning with PAC-Bayes bounds.
In recent years, deep reinforcement learning has been shown to be adept at solving sequential decision processes with high-dimensional state spaces such as in the Atari games. Many reinforcement learning problems, however, involve high-dimensional discrete action spaces as well as high-dimensional state spaces. This pa…
SPREV simplifies visualization of complex labeled datasets.
ADIGen: Automatic, Debiased, and Invariant Counterfactual Generation
CriticSMC improves planning efficiency in constrained environments.
In this paper, we present our approach to solve a physics-based reinforcement learning challenge "Learning to Run" with objective to train physiologically-based human model to navigate a complex obstacle course as quickly as possible. The environment is computationally expensive, has a high-dimensional continuous actio…
In many real-world scenarios, rewards extrinsic to the agent are extremely sparse, or absent altogether. In such cases, curiosity can serve as an intrinsic reward signal to enable the agent to explore its environment and learn skills that might be useful later in its life. We formulate curiosity as the error in an agen…
This paper considers improved forecasting in possibly nonlinear dynamic settings, with high-dimension predictors ("big data" environments). To overcome the curse of dimensionality and manage data and model complexity, we examine shrinkage estimation of a back-propagation algorithm of a deep neural net with skip-layer c…
Modeling agent behavior is central to understanding the emergence of complex phenomena in multiagent systems. Prior work in agent modeling has largely been task-specific and driven by hand-engineering domain-specific prior knowledge. We propose a general learning framework for modeling agent behavior in any multiagent …
Minimum attention improves reinforcement learning performance in high-dimensional dynamics.
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
A fundamental issue in reinforcement learning algorithms is the balance between exploration of the environment and exploitation of information already obtained by the agent. Especially, exploration has played a critical role for both efficiency and efficacy of the learning process. However, Existing works for explorati…
Efficient exploration is an unsolved problem in Reinforcement Learning which is usually addressed by reactively rewarding the agent for fortuitously encountering novel situations. This paper introduces an efficient active exploration algorithm, Model-Based Active eXploration (MAX), which uses an ensemble of forward mod…
Existing model-based reinforcement learning methods often study perception modeling and decision making separately. We introduce joint Perception and Control as Inference (PCI), a general framework to combine perception and control for partially observable environments through Bayesian inference. Based on the fact that…
Paper uses deep reinforcement learning for optimal stock portfolio management.
Modern reinforcement learning algorithms reach super-human performance on many board and video games, but they are sample inefficient, i.e. they typically require significantly more playing experience than humans to reach an equal performance level. To improve sample efficiency, an agent may build a model of the enviro…