Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

180360540720 · Jun 202019922001200920172026
48 results for Factored action space

FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.

problem Cooperative multi-agent reinforcement learning in discrete and continuous action spaces.
method FACMAC uses a centralised but factored critic, combining per-agent utilities into a joint action-value function.
result FACMAC outperforms MADDPG and other baselines on multi-agent particle environments and StarCraft II tasks.

Identifies latent actions and dynamics from offline data with diverse demonstrators.

problem Recovering latent actions and environment dynamics from action-free trajectories.
method Assumes distinct policies for each demonstrator, identifies latent transitions and policies via matrix factorization.
result Identifies latent transitions and demonstrator policies up to permutation.

Generalizing a classical theorem of Carlson and Toledo, we prove that any Zariski dense isometric action of a Kähler group on the real hyperbolic space of dimension at least 3 factors through a homomorphism onto a cocompact discrete subgroup of PSL(2,R). We also study actions of Kähler groups on infinite dimensional re…

2010-12-07abs ↗pdf ↗

We show that the Gromov boundary of the free factor graph for the free group Fn with n>2 generators is the space of equivalence classes of minimal very small indecomposable projective Fn-trees without point stabilizer containing a free factor equipped with a quotient topology. Here two such trees are equivalent if the …

2012-11-07abs ↗pdf ↗

We prove a multiplicity formula for Riemann-Roch numbers of reductions of Hamiltonian actions of loop groups. This includes as a special case the factorization formula for the quantum dimension of the moduli space of flat connections over a Riemann surface.

1996-12-30abs ↗pdf ↗

We study affine maps between CAT(0) spaces with geometric actions, and show that they essentially split as products of dilations and linear maps (on the Euclidean factor). This extends known results from the Riemannian case. Furthermore, we prove a splitting lemma for the Tits boundary of a CAT(0) space with geometric …

2013-09-04abs ↗pdf ↗

In this paper we prove that a fully irreducible outer automorphism relative to a non-exceptional free factor system acts loxodromically on the relative free factor complex as defined by Handel and Mosher. We also prove a north-south dynamic result for the action of such outer automorphisms on the closure of relative ou…

2016-12-13abs ↗pdf ↗

Study proper actions of Lie groups on symmetric spaces, finding rigidity results and Hurwitz-Radon numbers.

problem Proper actions of non-compact semisimple Lie groups on pseudo-Riemannian symmetric spaces.
method Analysis of symmetric spaces and rigidity results.
result Any connected non-compact semisimple Lie group acting properly on these spaces must be globally isomorphic to Spin(n,1)Spin(n,1) up to compact factors.

We characterize strongly Morse quasi-geodesics in Outer space as quasi-geodesics which project to quasi-geodesics in the free factor graph. We define convex cocompact subgroups of Out(Fn)Out(F_n) as subgroups such that an orbit map in the free factor graph is a quasi-isometric embedding, and we characterize such groups via …

2014-11-09abs ↗pdf ↗

Any continuous action of SL(n,Z), where n > 2, on a r-dimensional mod 2 homology sphere factors through a finite group action if r < n - 1. In particular, any continuous action of SL(n+2,Z) on the n-dimensional sphere factors through a finite group action.

2005-04-10abs ↗pdf ↗

A new RL paradigm reduces state-action-value function approximation inefficiency.

problem Challenges in state-action-value function approximation for RL.
method State Action Separable Reinforcement Learning (sasRL) decouples action space from value function learning.
result sasRL achieves up to 75% better performance than state-of-the-art MDP-based RL algorithms.

We show that an isometric action of a compact quantum group on the underlying geodesic metric space of a compact connected Riemannian manifold (M,g)(M,g) with strictly negative curvature is automatically classical, in the sense that it factors through the action of the isometry group of (M,g)(M,g). This partially answers a q…

2015-03-27abs ↗pdf ↗

In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…

2019-06-24abs ↗pdf ↗

Representation learning is a central challenge across a range of machine learning areas. In reinforcement learning, effective and functional representations have the potential to tremendously accelerate learning progress and solve more challenging problems. Most prior work on representation learning has focused on gene…

2018-11-19abs ↗pdf ↗

The study of pseudo-Anosov maps with minimum expansion factor using train tracks.

problem Finding pseudo-Anosov maps with minimum expansion factor.
method Analysis of standardly embedded train tracks and Thurston symplectic form.
result The expansion factor of pseudo-Anosov maps is bounded by a specific inequality involving the golden ratio.

New method for evaluating and learning in complex decision-making scenarios.

problem Evaluating and learning from policies in contextual combinatorial bandits with high bias and variance.
method Factored action space decomposition and importance sampling-based estimator (OPCB).
result OPCB achieves superior performance in OPE and OPL compared to conventional methods.

The goal of this article is to investigate nontrivial mm-quasi-Einstein manifolds globally conformal to an nn-dimensional Euclidean space. By considering such manifolds, whose conformal factors and potential functions are invariant under the action of an (n1)(n-1)-dimensional translation group, we provide a complete cl…

2019-04-25abs ↗pdf ↗

FPGs use structure to improve policy learning in complex tasks.

problem Policy gradient methods struggle with high-dimensional action spaces and objective multiplicity.
method Factor baseline and action-target influence network to reduce gradient variance.
result FPGs provide a general framework for state-of-the-art algorithms and improve performance.

We introduce reinforcement learning for heterogeneous teams in which rewards for an agent are additively factored into local costs, stimuli unique to each agent, and global rewards, those shared by all agents in the domain. Motivating domains include coordination of varied robotic platforms, which incur different costs…

2018-05-23abs ↗pdf ↗

We construct a local action of the group of rational maps from S2S^2 to GL(n,C)GL(n,C) on local solutions of flows of the ZS-AKNS sl(n,C)sl(n,C)-hierarchy. We show that the actions of simple elements (linear fractional transformations) give local Bäcklund transformations, and we derive a permutability formula from different fact…

1998-05-18abs ↗pdf ↗

We define a fuchsian affine action of a surface group to be such that the linear part factors through a representation of SL(2,R)SL(2,{\mathbb R}). We prove a fuchsian affine action of a surface group is never proper.

2000-05-25abs ↗pdf ↗

By means of a Kaluza-Klein type argument we show that the Perelman's F-functional is the Einstein-Hilbert action in a space with extra ``phantom'' dimensions. In this way, we try to interpret some remarks of Perelman in the introduction and at the end of the first section in his first famous paper. As a consequence the…

2008-05-21abs ↗pdf ↗

Study shows conditions for continuity of foliated homeomorphisms action on space of leaves.

problem Conditions for continuity of foliated homeomorphisms action on space of leaves.
method Investigated sufficient conditions for continuity of the homomorphism ψ: H(X, Δ) → H(Y) induced by the action of foliated homeomorphisms on the space of leaves.
result Similar results hold for a more general class of partitions of locally compact Hausdorff spaces.

Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model. We propose a parametric Q-learning algorithm that finds an approximate-optimal policy using a sample size proportional to the feature dimension KK and invariant …

2019-02-13abs ↗pdf ↗

We study the construction of quasimorphisms on groups acting on trees introduced by Monod and Shalom, that we call median quasimorphisms, and in particular we fully characterise actions on trees that give rise to non-trivial median quasimorphisms. Roughly speaking, either the action is highly transitive on geodesics, i…

2014-11-28abs ↗pdf ↗

We introduce the factored bandits model, which is a framework for learning with limited (bandit) feedback, where actions can be decomposed into a Cartesian product of atomic actions. Factored bandits incorporate rank-1 bandits as a special case, but significantly relax the assumptions on the form of the reward function…

2018-07-04abs ↗pdf ↗

Conditions for hyperbolic and relatively hyperbolic extensions of free groups using automorphisms with fixed points.

problem Conditions for hyperbolic and relatively hyperbolic extensions of free groups.
method Using dynamics of outer automorphisms on the complex of free factors and investigating the geometry of the extension group.
result Conditions for hyperbolic and relatively hyperbolic extensions of free groups using automorphisms with fixed points.

Improved exploration in factored average-reward MDPs reduces regret.

problem Minimizing regret in unknown Factored Markov Decision Processes (FMDPs).
method DBN-UCRL strategy, inspired by UCRL2, uses Bernstein-type confidence sets for individual elements of the transition function.
result Achieves a regret bound with a leading term strictly improving over existing bounds.

Any reinforcement learning algorithm that applies to all Markov decision processes (MDPs) will suffer Ω(SAT)Ω(\sqrt{SAT}) regret on some MDP, where TT is the elapsed time and SS and AA are the cardinalities of the state and action spaces. This implies T=Ω(SA)T = Ω(SA) time to guarantee a near-optimal policy. In many settings…

2014-03-15abs ↗pdf ↗

We consider interactive learning and covering problems, in a setting where actions may incur different costs, depending on the response to the action. We propose a natural greedy algorithm for response-dependent costs. We bound the approximation factor of this greedy algorithm in active learning settings as well as in …

2016-02-23abs ↗pdf ↗