A new RL paradigm reduces state-action-value function approximation inefficiency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In recent times, the use of separable convolutions in deep convolutional neural network architectures has been explored. Several researchers, most notably (Chollet, 2016) and (Ghosh, 2017) have used separable convolutions in their deep architectures and have demonstrated state of the art or close to state of the art pe…
Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actions in the demonstrations to be fully available, which is hard to ensure in real applications. Though algorithms for learning with unobservab…
Network slicing promises to provision diversified services with distinct requirements in one infrastructure. Deep reinforcement learning (e.g., deep -learning, DQL) is assumed to be an appropriate algorithm to solve the demand-aware inter-slice resource management issue in network slicing by regarding the …
Characterizes geometric actions on graphs with flexible stabilizers.
This paper proposes a simple yet effective method for human action recognition in video. The proposed method separately extracts local appearance and motion features using state-of-the-art three-dimensional convolutional neural networks from sampled snippets of a video. These local features are then concatenated to for…
Fast covariance calculation is required both for SLAM (e.g.~in order to solve data association) and for evaluating the information-theoretic term for different candidate actions in belief space planning (BSP). In this paper we make two primary contributions. First, we develop a novel general-purpose incremental covaria…
This paper simplifies complex game dynamics by using a recursive representation.
Sharp conditions link separators to R-trees for space transformations.
A hyperkähler 4-metric with a triholomorphic SU(2) action gives rise to a family of confocal quadrics in Euclidean 3-space when cast in the canonical form of a hyperkähler 4-metric metric with a triholomorphic circle action. Moreover, at least in the case of geodesics orthogonal to the U(1) fibres, both the covariant S…
The paper builds complex hyperbolic 2-manifolds with isolated singularities.
In a previous article, analytic 1-submanifolds had been classified w.r.t. their symmetry under a given regular and separately analytic Lie group action on an analytic manifold. It was shown that such an analytic 1-submanifold is either free or (via the exponential map) analytically diffeomorphic to the unit circle or a…
Proposes a value-based method for continuous control without an actor.
Exploration is an extremely challenging problem in reinforcement learning, especially in high dimensional state and action spaces and when only sparse rewards are available. Effective representations can indicate which components of the state are task relevant and thus reduce the dimensionality of the space to explore.…
Q()-Learning improves Q-Learning by separating action-value functions into different time scales.
Building agents to interact with the web would allow for significant improvements in knowledge understanding and representation learning. However, web navigation tasks are difficult for current deep reinforcement learning (RL) models due to the large discrete action space and the varying number of actions between the s…
We use the theory of group actions on profinite trees to prove that the fundamental group of a finite, 1-acylindrical graph of free groups with finitely generated edge groups is conjugacy separable. This has several applications: we prove that positive, one-relator groups are conjugacy separable; we provide a…
Let M be a hyperbolizable, nontrivial compression body without toroidal boundary components. In this paper, we characterize which discrete and faithful representations of the fundamental group of M into PSL(2,C) are separable-stable. The set of separable-stable representations forms a domain of discontinuity for the ac…
AI learns to design chemical processes efficiently.
We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…
We construct an example of an isometric action of on a -hyperbolic graph , such that this action is acylindrical, purely loxodromic, has asymptotic translation lengths of nontrivial elements of separated away from , has quasiconvex orbits in , but such that the orbit map is n…
One-shot path planning for multiple agents using neural networks.
Optimistic initialisation is an effective strategy for efficient exploration in reinforcement learning (RL). In the tabular case, all provably efficient model-free algorithms rely on it. However, model-free deep RL algorithms do not use optimistic initialisation despite taking inspiration from these provably efficient …
Study shows how to learn optimal policies quickly in stochastic control problems.
The growing use of virtual autonomous agents in applications like games and entertainment demands better control policies for natural-looking movements and actions. Unlike the conventional approach of hard-coding motion routines, we propose a deep learning method for obtaining control policies by directly mimicking raw…
We prove that the set of orthogonal separable coordinates on an arbitrary (pseudo-)Riemannian manifold carries a natural structure of a projective variety, equipped with an action of the isometry group. This leads us to propose a new, algebraic geometric approach to the classification of orthogonal separable coordinate…
The interplay between the Hamilton-Jacobi theory of orthogonal separation of variables and the theory of group actions is investigated based on concrete examples.
No exceptional orbits found in Hilbert spaces actions.
This paper presents a new meta-modeling framework to employ deep reinforcement learning (DRL) to generate mechanical constitutive models for interfaces. The constitutive models are conceptualized as information flow in directed graphs. The process of writing constitutive models are simplified as a sequence of forming g…
New complex connects graph separability to group properties.
Quantum states associated with subsets of product manifolds are separable.
The paper proposes a method to identify power system oscillation modes using blind source separation.
New algorithm for reward-free RL with linear function approximation, reducing sample complexity.
The study quantifies how many objects can be linearly classified under all views.
This paper sets a lower bound for sample complexity in inverse reinforcement learning.
New machine learning method detects quantum separability in large-scale systems.
New reinforcement learning framework for adapting to new actions.
Paper learns meaningful state and action representations from MDP trajectories.
Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot scale to tasks with very high state and action dimensionality such as 3D humanoid l…
Recent work has shown that reinforcement learning (RL) is a promising approach to control dynamical systems described by partial differential equations (PDE). This paper shows how to use RL to tackle more general PDE control problems that have continuous high-dimensional action spaces with spatial relationship among ac…
Analytic curves are classified w.r.t. their symmetry under a regular and separately analytic Lie group action on an analytic manifold. We show that an analytic curve is either exponential or splits into countably many analytic immersive curves, each of them discretely generated by the symmetry group (i.e., each such cu…
We compute the quotient of the self-duality equation for conformal metrics by the action of the diffeomorphism group. We also determine Hilbert polynomial, counting the number of independent scalar differential invariants depending on the jet-order, and the corresponding Poincaré function. We describe the field of rati…
QTRAN++ improves MARL performance in complex environments.
TensorPlan shows an exponential lower bound for planning in MDPs with linearly realizable value functions.
This paper solves the normalizability crisis in sequential inference by introducing bounded information geometry.
Calculates Dehn twist actions on conformal blocks for modular categories.
Paper examines Dehn twists on non-orientable surfaces and their limitations.
New RL method learns from state transitions without actions.