Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

22446688 · Jun 202019922001200920172026
48 results for Concurrent RL

Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.

problem Learning more data-efficient and better policies in concurrent RL with non-identical environments.
method Proposes a novel algorithmic framework that leverages causal inference via ANM-MM to extract model parameters and a new data sharing scheme based on similarity measures.
result Demonstrates superior learning speeds on various tasks and effectiveness of diverse action selection.

In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …

2018-03-14abs ↗pdf ↗

New offline RL method works with limited data and function approximators.

problem Sample efficiency with limited data and weak function approximators.
method Pessimistic algorithm based on version space formed by marginalized importance sampling (MIS), with gap assumption.
result Guarantees sample efficiency for simple algorithm under specific assumptions.

Meta-learning is a tool that allows us to build sample-efficient learning systems. Here we show that, once meta-trained, LSTM Meta-Learners aren't just faster learners than their sample-inefficient deep learning (DL) and reinforcement learning (RL) brethren, but that they actually pursue fundamentally different learnin…

2019-05-03abs ↗pdf ↗

In many real-world reinforcement learning (RL) problems, besides optimizing the main objective function, an agent must concurrently avoid violating a number of constraints. In particular, besides optimizing performance it is crucial to guarantee the safety of an agent during training as well as deployment (e.g. a robot…

2018-05-20abs ↗pdf ↗

This paper examines a generalized Kropina metric and its geometric properties.

problem Investigating geometric properties of a generalized Kropina metric.
method Analyzing a Finsler manifold with a concurrent π-vector field and examining the φφ-concurrent generalized Kropina change.
result The geodesic sprays of the original Finsler metric and the modified metric are never projectively related.

The paper studies Finsler spaces with semi-concurrent vector fields and their equivalence to Riemannian spaces.

problem Characterizing Finsler spaces with semi-concurrent vector fields.
method Analyzing various Finsler spaces and proving conditions for equivalence to Riemannian spaces.
result Various Finsler spaces (quasi-CC-reducible, C3C3-like, ChC^{h}-recurrent, P2P2-like) are equivalent to Riemannian spaces if they admit a semi-concurrent vector field.

A Ricci soliton (Mn,g,v,λ)(M^n,g,v,λ) on a Riemannian manifold (Mn,g)(M^n,g) is said to have concurrent potential field if its potential field vv is a concurrent vector field. In the first part of this paper we completely classify Ricci solitons with concurrent potential fields. In the second part we derive a necessary and suffic…

2014-07-10abs ↗pdf ↗

A Ricci soliton (M,g,v,λ)(M,g,v,λ) on a Riemannian manifold (M,g)(M,g) is said to have concurrent potential field if its potential field vv is a concurrent vector field. Ricci solitons arisen from concurrent vector fields on Riemannian manifolds were studied recently in \cite{CD2}. The most important concurrent vector field is …

2014-10-19abs ↗pdf ↗

In the present paper, we introduce and investigate the notion of a semi concurrent vector field on a Finsler manifold. We show that some special Finsler manifolds admitting such vector fields turn out to be Riemannian. We prove that Tachibana's characterization of Finsler manifolds admitting a concurrent vector field l…

2018-02-07abs ↗pdf ↗

Paper applies RL to optimize inventory management across multiple products and nodes.

problem Optimizing inventory management for a large number of products with shared capacity in a multi-node supply chain.
method Novel multi-agent hierarchical reinforcement learning framework with A2C algorithm and quantised action spaces.
result The approach optimizes for maximizing product sales and minimizing wastage of perishable products.

This research tackles balancing exploration and exploitation in deep RL for partially observable systems.

problem Balancing exploration and exploitation in deep RL for partially observable systems.
method Deployed and tested several techniques including adaptive and deterministic exploration strategies, and a modified quadratic loss function.
result Adaptive methods better approximate the trade-off between exploration and exploitation.

A new reinforcement learning method for robots thinking and moving simultaneously.

problem Concurrent control in robotic systems where actions must be decided while the system is still evolving.
method Continuous-time Bellman equations, discretization aware of system delays, and architectural extension to deep reinforcement learning.
result The method successfully handles tasks requiring simultaneous decision-making and action execution.

BCO* improves BCO by concurrently training inverse dynamics and expert policy.

problem Efficiently learn from unlabeled demonstrations without requiring many initial interactions.
method Introduce BCO* that concurrently trains an inverse dynamics model and expert policy.
result BCO* eliminates the need for initial interactions and improves sample complexity.

The present paper deals with an \emph{intrinsic} investigation of the notion of a concurrent ππ-vector field on the pullback bundle of a Finsler manifold (M,L)(M,L). The effect of the existence of a concurrent ππ-vector field on some important special Finsler spaces is studied. An intrinsic investigation of a particular…

2008-05-16abs ↗pdf ↗

We consider the problem of concurrent portfolio losses in two non-overlapping credit portfolios. In order to explore the full statistical dependence structure of such portfolio losses, we estimate their empirical pairwise copulas. Instead of a Gaussian dependence, we typically find a strong asymmetry in the copulas. Co…

2016-04-23abs ↗pdf ↗

Paper proposes a method to learn and exceed expert demonstrations in unknown reward environments.

problem Learning to outperform expert demonstrations in unknown reward environments.
method A novel concurrent reward and action policy learning approach with a stereo utility definition.
result The proposed method can outperform expert demonstrations in various environments.

In this paper, we completely classify almost Yamabe solitons on hypersurfaces in Euclidean spaces arisen from the position vector field. Some results of almost Yamabe solitons with a concurrent vector field and almost Yamabe solitons on submanifolds in Riemannian manifolds equipped with a concurrent vector field are al…

2017-11-13abs ↗pdf ↗

We generalize Matsumoto metrics with a special π-form and explore their geometric properties.

problem Exploring the geometric properties of generalized Matsumoto metrics with a special π-form.
method Considering a Finsler manifold with a concurrent π-vector field, we introduce a change in the metric and analyze its geometric properties.
result The generalized φ-Matsumoto metric can never be projectively related to the original metric.

Deep learning model classifies concurrent human interactions from WiFi data with high accuracy.

problem Classifying concurrent human interactions from WiFi data with high accuracy.
method Attention-BiGRU deep learning model using Multiple Input Multiple Output radio link.
result Maximum benchmark accuracy of 94% for a single subject-pair, 88% for ten subject pairs.

Improved gap-dependent bounds for reinforcement learning with linear approximations.

problem Achieving nearly minimax-optimal performance with linear function approximation.
method Developed and analyzed the LSVI-UCB++ algorithm and its concurrent variant.
result First gap-dependent regret bound for nearly minimax-optimal algorithm LSVI-UCB++.

New algorithm trains neural nets on simple skills to learn complex tasks faster.

problem Learning complex tasks through simple imitation.
method Train neural networks on simple, easy-to-learn skills to accelerate learning of complex, hard-to-learn tasks.
result Consistently outperforms state-of-the-art baseline in training speed and performance.

A novel method optimizes variable-stiffness structures for better strength and weight.

problem Optimizing variable-stiffness structures for higher strength and lighter weight.
method A novel multi-stage concurrent topology optimization scheme combining DMO, S-BPTO, and CFAO.
result The method ensures better fibre angle convergence and stable optimization.

We study the equilibrium positions of three points on a convex curve under influence of the Coulomb potential. We identify these positions as orthotripods, three points on the curve having concurrent normals. This relates the equilibrium positions to the caustic (evolute) of the curve. The concurrent normals can only m…

2015-03-14abs ↗pdf ↗

We consider a team of reinforcement learning agents that concurrently operate in a common environment, and we develop an approach to efficient coordinated exploration that is suitable for problems of practical scale. Our approach builds on seed sampling (Dimakopoulou and Van Roy, 2018) and randomized value function lea…

2018-05-23abs ↗pdf ↗

The study confirms conjectures about normals to convex polytopes in 3D space.

problem Concurrent normals problem for convex polytopes in 3D.
method Analyzes the PL concurrent normals problem for convex polytopes, proving conjectures for specific cases.
result Polytopes in 3D have points with 10 normals from interior points, confirmed for all tetrahedra and triangular prisms.

The paper discusses the impossibility of eliminating surplus intersections in Lagrangian submanifolds.

problem Can surplus intersections in Lagrangian submanifolds be eliminated by Hamiltonian isotopy?
method Analyzing the intersections and isotopies of Lagrangian submanifolds and auxiliary Lagrangians.
result Surplusection cannot be eliminated in several important situations, highlighting the need for better understanding.

This work explores representation complexity in RL paradigms, revealing model-based RL as the easiest task.

problem Investigating the representation complexity gap among model-based, policy-based, and value-based RL.
method Demonstrated through analysis of Markov decision processes (MDPs) and introduced new classes of MDPs.
result Representation complexity hierarchy: model-based RL > policy-based RL > value-based RL.

Study on vector fields on Lie groups reveals surprising algebraic coincidences.

problem Characterizing vector fields on Lie groups with Riemannian metrics.
method Algebraic and geometric analysis of left-invariant vector fields on nilpotent Lie groups.
result Spaces of Killing, one-harmonic, and conformal vector fields coincide with the center of the Lie algebra on nilpotent Lie groups.

New method defends RL agents from poisoning attacks without MDP knowledge.

problem Poisoning attacks on RL systems can cause learning failures.
method Generic poisoning framework for online RL, Vulnerability-Aware Adversarial Critic Poison (VA2C-P).
result Successfully prevents RL agents from learning good policies or converging to target policies.

In dialogues, an utterance is a chain of consecutive sentences produced by one speaker which ranges from a short sentence to a thousand-word post. When studying dialogues at the utterance level, it is not uncommon that an utterance would serve multiple functions. For instance, "Thank you. It works great." expresses bot…

2019-09-02abs ↗pdf ↗

Reincarnating RL reuses prior work to accelerate RL progress.

problem Efficiency and accessibility in reinforcement learning for large-scale applications.
method Transfer of learned policies between RL agents or design iterations, focusing on value-based RL.
result Demonstrated gains in performance over tabula rasa RL on various tasks.

RL tackles decision making in unknown environments, focusing on efficiency and efficacy.

problem Efficiency and efficacy in RL algorithms for sample-starved situations.
method Markov Decision Processes, model-based and value-based approaches, policy optimization.
result Enhanced understanding and improvements in sample and computational efficacies of RL algorithms.

Study shows effectiveness of offline RL in online RL tasks.

problem Improving online RL efficiency using offline RL data.
method Formalized framework for incorporating offline RL as online RL subroutines, introducing techniques to enhance effectiveness.
result Effectiveness of the framework depends on task nature, techniques greatly enhance effectiveness, and existing methods are ineffective.