Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

181361542722 · Jun 202019922001200920172026
48 results for State Action Separable

A new RL paradigm reduces state-action-value function approximation inefficiency.

problem Challenges in state-action-value function approximation for RL.
method State Action Separable Reinforcement Learning (sasRL) decouples action space from value function learning.
result sasRL achieves up to 75% better performance than state-of-the-art MDP-based RL algorithms.

In recent times, the use of separable convolutions in deep convolutional neural network architectures has been explored. Several researchers, most notably (Chollet, 2016) and (Ghosh, 2017) have used separable convolutions in their deep architectures and have demonstrated state of the art or close to state of the art pe…

2017-01-16abs ↗pdf ↗

Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actions in the demonstrations to be fully available, which is hard to ensure in real applications. Though algorithms for learning with unobservab…

2019-05-29abs ↗pdf ↗

Characterizes geometric actions on graphs with flexible stabilizers.

problem Understanding geometric actions on flexible stabilizers.
method Defining generalized fine actions and proving relative quasi-convexity criteria.
result Characterizes Bowditch boundary points in relatively geometric actions.

The paper builds complex hyperbolic 2-manifolds with isolated singularities.

problem Finding compact complex hyperbolic 2-manifolds with non-free actions and isolated fixed points.
method Constructs specific examples for each prime pp with Z/pZ\mathbb{Z} / p \mathbb{Z} action.
result General examples for p=2p=2 related to complex hyperbolic lattices conjugacy separability.

In a previous article, analytic 1-submanifolds had been classified w.r.t. their symmetry under a given regular and separately analytic Lie group action on an analytic manifold. It was shown that such an analytic 1-submanifold is either free or (via the exponential map) analytically diffeomorphic to the unit circle or a…

2016-01-26abs ↗pdf ↗

Exploration is an extremely challenging problem in reinforcement learning, especially in high dimensional state and action spaces and when only sparse rewards are available. Effective representations can indicate which components of the state are task relevant and thus reduce the dimensionality of the space to explore.…

2019-05-27abs ↗pdf ↗

Q(ΔΔ)-Learning improves Q-Learning by separating action-value functions into different time scales.

problem Q-Learning struggles with bias-variance trade-off, especially in long-term rewards.
method Introduces Q(ΔΔ)-Learning, extending TD(ΔΔ) to decompose Q(ΔΔ)-function into distinct discount factors.
result Q(ΔΔ)-Learning achieves better stability and scalability, especially for long-term tasks.

Building agents to interact with the web would allow for significant improvements in knowledge understanding and representation learning. However, web navigation tasks are difficult for current deep reinforcement learning (RL) models due to the large discrete action space and the varying number of actions between the s…

2019-02-19abs ↗pdf ↗

We use the theory of group actions on profinite trees to prove that the fundamental group of a finite, 1-acylindrical graph of free groups with finitely generated edge groups is conjugacy separable. This has several applications: we prove that positive, C(1/6)C'(1/6) one-relator groups are conjugacy separable; we provide a…

2009-05-30abs ↗pdf ↗

Let M be a hyperbolizable, nontrivial compression body without toroidal boundary components. In this paper, we characterize which discrete and faithful representations of the fundamental group of M into PSL(2,C) are separable-stable. The set of separable-stable representations forms a domain of discontinuity for the ac…

2013-11-06abs ↗pdf ↗

We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…

2017-02-28abs ↗pdf ↗

We construct an example of an isometric action of F(a,b)F(a,b) on a δδ-hyperbolic graph YY, such that this action is acylindrical, purely loxodromic, has asymptotic translation lengths of nontrivial elements of F(a,b)F(a,b) separated away from 00, has quasiconvex orbits in YY, but such that the orbit map F(a,b)YF(a,b)\to Y is n…

2015-04-20abs ↗pdf ↗

One-shot path planning for multiple agents using neural networks.

problem Efficiently generating optimal or near-optimal paths for multiple agents in robotics.
method Utilizes fully convolutional neural networks for one-shot multi-agent path planning.
result Demonstrates successful generation of optimal or near-optimal paths in over 85% of cases for multi-path planning.

Optimistic initialisation is an effective strategy for efficient exploration in reinforcement learning (RL). In the tabular case, all provably efficient model-free algorithms rely on it. However, model-free deep RL algorithms do not use optimistic initialisation despite taking inspiration from these provably efficient …

2020-02-26abs ↗pdf ↗

Study shows how to learn optimal policies quickly in stochastic control problems.

problem Learning optimal policies in large, continuous state and action spaces with limited data.
method Analyzes three geometric exponents to quantify fast policy regret convergence.
result Shows that fast policy regret convergence is induced by specific geometric structures.

We prove that the set of orthogonal separable coordinates on an arbitrary (pseudo-)Riemannian manifold carries a natural structure of a projective variety, equipped with an action of the isometry group. This leads us to propose a new, algebraic geometric approach to the classification of orthogonal separable coordinate…

2015-10-30abs ↗pdf ↗

The paper proposes a method to identify power system oscillation modes using blind source separation.

problem Accurately identifying oscillation modes in power systems with renewable energy sources.
method A high-order blind source identification (HOBI) algorithm based on copula statistic combined with Hilbert transform and iteration procedure.
result The method can identify all oscillation modes and model order from a single channel of observation signals, outperforming state-of-the-art methods.

New algorithm for reward-free RL with linear function approximation, reducing sample complexity.

problem Efficiently learning optimal policies without prior reward information in complex environments.
method Developed an algorithm for reward-free RL in linear Markov decision processes, proving sample complexity bounds.
result Polynomial sample complexity in feature dimension and planning horizon, independent of states and actions.

The study quantifies how many objects can be linearly classified under all views.

problem Understanding the expressivity of group-equivariant representations.
method Generalization of Cover's Function Counting Theorem to quantify separable dichotomies.
result The fraction of separable dichotomies is determined by the fixed space dimension of the group action.

This paper sets a lower bound for sample complexity in inverse reinforcement learning.

problem Finding a reward function that generates a desired optimal policy in MDPs.
method Information-theoretic lower bound using geometric construction and Fano's inequality.
result An O(nlogn)O(n \log n) sample complexity lower bound for IRL problems.

New machine learning method detects quantum separability in large-scale systems.

problem Deciding quantum separability of large-scale bipartite density matrices.
method Frank-Wolfe-based algorithm for finding nearest separable density matrices and classification of density matrices as separable or entangled.
result The method scales up to thousands of density matrices and achieves high quantum entanglement detection accuracy.

Paper learns meaningful state and action representations from MDP trajectories.

problem Learning good state and action representations from MDP trajectories.
method Tensor decomposition, kernelization, importance sampling, low-Tucker-rank approximation.
result The learned state/action abstractions provide accurate approximations to latent block structures.

Analytic curves are classified w.r.t. their symmetry under a regular and separately analytic Lie group action on an analytic manifold. We show that an analytic curve is either exponential or splits into countably many analytic immersive curves, each of them discretely generated by the symmetry group (i.e., each such cu…

2016-01-25abs ↗pdf ↗

We compute the quotient of the self-duality equation for conformal metrics by the action of the diffeomorphism group. We also determine Hilbert polynomial, counting the number of independent scalar differential invariants depending on the jet-order, and the corresponding Poincaré function. We describe the field of rati…

2016-05-04abs ↗pdf ↗

QTRAN++ improves MARL performance in complex environments.

problem Poor empirical performance of QTRAN in complex environments.
method Stabilizing training objective, removing role separation, and introducing a multi-head mixing network.
result QTRAN++ achieves state-of-the-art performance in the Starcraft Multi-Agent Challenge (SMAC).

TensorPlan shows an exponential lower bound for planning in MDPs with linearly realizable value functions.

problem Finding an exponential lower bound for planning in MDPs with linearly realizable value functions.
method TensorPlan and a few action lower bound approach.
result An exponentially large lower bound is shown for planning in MDPs with linearly realizable value functions.

This paper solves the normalizability crisis in sequential inference by introducing bounded information geometry.

problem Structural failure in standard sequential inference architectures when dealing with extreme outliers.
method Non-parametric field actions and bounded information geometry to truncate infinite tails of spatial distributions.
result Empirical benchmarks across three domains show robust estimation without infinite-tailed distributional assumptions.

Paper examines Dehn twists on non-orientable surfaces and their limitations.

problem Limitations of generating Dehn twists on non-orientable surfaces.
method Analyzes the level 2 mapping class group of non-orientable surfaces and their subgroups.
result Dehn twist subgroup of M2(Ng)\mathcal{M}_2(N_g) cannot be generated by squares of Dehn twists about non-separating curves.