Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …
New offline RL method works with limited data and function approximators.
Meta-learning is a tool that allows us to build sample-efficient learning systems. Here we show that, once meta-trained, LSTM Meta-Learners aren't just faster learners than their sample-inefficient deep learning (DL) and reinforcement learning (RL) brethren, but that they actually pursue fundamentally different learnin…
A new method improves ridesharing efficiency using QMIX.
In many real-world reinforcement learning (RL) problems, besides optimizing the main objective function, an agent must concurrently avoid violating a number of constraints. In particular, besides optimizing performance it is crucial to guarantee the safety of an agent during training as well as deployment (e.g. a robot…
This paper examines a generalized Kropina metric and its geometric properties.
New approach uses unlabeled prior data to accelerate exploration in sparse reward tasks.
The paper studies Finsler spaces with semi-concurrent vector fields and their equivalence to Riemannian spaces.
A Ricci soliton on a Riemannian manifold is said to have concurrent potential field if its potential field is a concurrent vector field. In the first part of this paper we completely classify Ricci solitons with concurrent potential fields. In the second part we derive a necessary and suffic…
A Ricci soliton on a Riemannian manifold is said to have concurrent potential field if its potential field is a concurrent vector field. Ricci solitons arisen from concurrent vector fields on Riemannian manifolds were studied recently in \cite{CD2}. The most important concurrent vector field is …
In the present paper, we introduce and investigate the notion of a semi concurrent vector field on a Finsler manifold. We show that some special Finsler manifolds admitting such vector fields turn out to be Riemannian. We prove that Tachibana's characterization of Finsler manifolds admitting a concurrent vector field l…
We propose a mechanism for distributed resource management and interference mitigation in wireless networks using multi-agent deep reinforcement learning (RL). We equip each transmitter in the network with a deep RL agent that receives delayed observations from its associated users, while also exchanging observations w…
Modelling and exploiting teammates' policies in cooperative multi-agent systems have long been an interest and also a big challenge for the reinforcement learning (RL) community. The interest lies in the fact that if the agent knows the teammates' policies, it can adjust its own policy accordingly to arrive at proper c…
Speeds up deep neural networks training by 10x using GPU concurrency.
Paper applies RL to optimize inventory management across multiple products and nodes.
This research tackles balancing exploration and exploitation in deep RL for partially observable systems.
A new reinforcement learning method for robots thinking and moving simultaneously.
BCO* improves BCO by concurrently training inverse dynamics and expert policy.
The present paper deals with an \emph{intrinsic} investigation of the notion of a concurrent -vector field on the pullback bundle of a Finsler manifold . The effect of the existence of a concurrent -vector field on some important special Finsler spaces is studied. An intrinsic investigation of a particular…
We consider the problem of concurrent portfolio losses in two non-overlapping credit portfolios. In order to explore the full statistical dependence structure of such portfolio losses, we estimate their empirical pairwise copulas. Instead of a Gaussian dependence, we typically find a strong asymmetry in the copulas. Co…
Paper proposes a method to learn and exceed expert demonstrations in unknown reward environments.
In this paper, we completely classify almost Yamabe solitons on hypersurfaces in Euclidean spaces arisen from the position vector field. Some results of almost Yamabe solitons with a concurrent vector field and almost Yamabe solitons on submanifolds in Riemannian manifolds equipped with a concurrent vector field are al…
Recent progress in artificial intelligence through reinforcement learning (RL) has shown great success on increasingly complex single-agent environments and two-player turn-based games. However, the real-world contains multiple agents, each learning and acting independently to cooperate and compete with other agents, a…
We generalize Matsumoto metrics with a special π-form and explore their geometric properties.
Non-negative matrix factorization (NMF) is a fundamental non-convex optimization problem with numerous applications in Machine Learning (music analysis, document clustering, speech-source separation etc). Despite having received extensive study, it is poorly understood whether or not there exist natural algorithms that…
Deep learning model classifies concurrent human interactions from WiFi data with high accuracy.
Improved gap-dependent bounds for reinforcement learning with linear approximations.
Training neural network often uses a machine learning framework such as TensorFlow and Caffe2. These frameworks employ a dataflow model where the NN training is modeled as a directed graph composed of a set of nodes. Operations in neural network training are typically implemented by the frameworks as primitives and rep…
New algorithm trains neural nets on simple skills to learn complex tasks faster.
A novel method optimizes variable-stiffness structures for better strength and weight.
A decentralized deep RL controller improves hexapod locomotion learning.
We study the equilibrium positions of three points on a convex curve under influence of the Coulomb potential. We identify these positions as orthotripods, three points on the curve having concurrent normals. This relates the equilibrium positions to the caustic (evolute) of the curve. The concurrent normals can only m…
We consider a team of reinforcement learning agents that concurrently operate in a common environment, and we develop an approach to efficient coordinated exploration that is suitable for problems of practical scale. Our approach builds on seed sampling (Dimakopoulou and Van Roy, 2018) and randomized value function lea…
Interference among concurrent transmissions in a wireless network is a key factor limiting the system performance. One way to alleviate this problem is to manage the radio resources in order to maximize either the average or the worst-case performance. However, joint consideration of both metrics is often neglected as …
The study confirms conjectures about normals to convex polytopes in 3D space.
Hybrid RL algorithms improve offline and online RL in linear MDPs.
The paper discusses the impossibility of eliminating surplus intersections in Lagrangian submanifolds.
This work explores representation complexity in RL paradigms, revealing model-based RL as the easiest task.
Study on vector fields on Lie groups reveals surprising algebraic coincidences.
The aim of this paper is to train an RBF neural network and select centers under concurrent faults. It is well known that fault tolerance is a very attractive property for neural networks. And center selection is an important procedure during the training process of an RBF neural network. In this paper, we devise two n…
New method defends RL agents from poisoning attacks without MDP knowledge.
In dialogues, an utterance is a chain of consecutive sentences produced by one speaker which ranges from a short sentence to a thousand-word post. When studying dialogues at the utterance level, it is not uncommon that an utterance would serve multiple functions. For instance, "Thank you. It works great." expresses bot…
Reincarnating RL reuses prior work to accelerate RL progress.
Paper shows RLHF can be solved similarly to standard RL.
RL tackles decision making in unknown environments, focusing on efficiency and efficacy.
Catalyst.RL accelerates RL research with efficient training.
Study shows effectiveness of offline RL in online RL tasks.