New method reduces bias and variance in OPE for large action spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
One problem in the application of reinforcement learning to real-world problems is the curse of dimensionality on the action space. Macro actions, a sequence of primitive actions, have been studied to diminish the dimensionality of the action space with regard to the time axis. However, previous studies relied on human…
FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.
In an effort to better understand the different ways in which the discount factor affects the optimization process in reinforcement learning, we designed a set of experiments to study each effect in isolation. Our analysis reveals that the common perception that poor performance of low discount factors is caused by (to…
New structural result classifies actions on symmetric spaces.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
Cohomogeneity-one actions on symmetric spaces of mixed type
Generalizing a classical theorem of Carlson and Toledo, we prove that any Zariski dense isometric action of a Kähler group on the real hyperbolic space of dimension at least 3 factors through a homomorphism onto a cocompact discrete subgroup of PSL(2,R). We also study actions of Kähler groups on infinite dimensional re…
We show that the Gromov boundary of the free factor graph for the free group Fn with n>2 generators is the space of equivalence classes of minimal very small indecomposable projective Fn-trees without point stabilizer containing a free factor equipped with a quotient topology. Here two such trees are equivalent if the …
We prove a multiplicity formula for Riemann-Roch numbers of reductions of Hamiltonian actions of loop groups. This includes as a special case the factorization formula for the quantum dimension of the moduli space of flat connections over a Riemann surface.
We study affine maps between CAT(0) spaces with geometric actions, and show that they essentially split as products of dilations and linear maps (on the Euclidean factor). This extends known results from the Riemannian case. Furthermore, we prove a splitting lemma for the Tits boundary of a CAT(0) space with geometric …
In this paper we prove that a fully irreducible outer automorphism relative to a non-exceptional free factor system acts loxodromically on the relative free factor complex as defined by Handel and Mosher. We also prove a north-south dynamic result for the action of such outer automorphisms on the closure of relative ou…
Study proper actions of Lie groups on symmetric spaces, finding rigidity results and Hurwitz-Radon numbers.
We characterize strongly Morse quasi-geodesics in Outer space as quasi-geodesics which project to quasi-geodesics in the free factor graph. We define convex cocompact subgroups of as subgroups such that an orbit map in the free factor graph is a quasi-isometric embedding, and we characterize such groups via …
Any continuous action of SL(n,Z), where n > 2, on a r-dimensional mod 2 homology sphere factors through a finite group action if r < n - 1. In particular, any continuous action of SL(n+2,Z) on the n-dimensional sphere factors through a finite group action.
A new RL paradigm reduces state-action-value function approximation inefficiency.
We show that an isometric action of a compact quantum group on the underlying geodesic metric space of a compact connected Riemannian manifold with strictly negative curvature is automatically classical, in the sense that it factors through the action of the isometry group of . This partially answers a q…
Decouples homotopy quotients of generalised configuration spaces on surfaces.
The aim of this paper is to study cohomogeneity one isometric linear actions on the -dimensional pseudo-Euclidean space . It is proved that the natural isometric action of the nilpotent factor of an Iwasawa decomposition of is not of cohomogeneity one. The orbits of cohomogeneity one ac…
Optimal control in latent factor models uses Tsallis entropy for exploration.
In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…
Representation learning is a central challenge across a range of machine learning areas. In reinforcement learning, effective and functional representations have the potential to tremendously accelerate learning progress and solve more challenging problems. Most prior work on representation learning has focused on gene…
The study of pseudo-Anosov maps with minimum expansion factor using train tracks.
New method for evaluating and learning in complex decision-making scenarios.
We prove that among all compact homogeneous spaces for an effective transitive action of a Lie group whose Levi subgroup has no compact simple factors, the seven-dimensional flat torus is the only one that admits an invariant torsion-free -structure.
Classifies cobounded hyperbolic actions of metabelian groups.
New topology shows Morse boundaries are topologically invariant.
The goal of this article is to investigate nontrivial -quasi-Einstein manifolds globally conformal to an -dimensional Euclidean space. By considering such manifolds, whose conformal factors and potential functions are invariant under the action of an -dimensional translation group, we provide a complete cl…
FPGs use structure to improve policy learning in complex tasks.
We prove that the exceptional complex Lie group has a transitive action on the hyperplane section of the complex Cayley plane . Our proof is direct and constructive. We use an explicit realization of the vector and spin actions of $\Spin(9,\C) \leq F_4$. Moreover, we identify the stabilizer of the …
We introduce reinforcement learning for heterogeneous teams in which rewards for an agent are additively factored into local costs, stimuli unique to each agent, and global rewards, those shared by all agents in the domain. Motivating domains include coordination of varied robotic platforms, which incur different costs…
We construct a local action of the group of rational maps from to on local solutions of flows of the ZS-AKNS -hierarchy. We show that the actions of simple elements (linear fractional transformations) give local Bäcklund transformations, and we derive a permutability formula from different fact…
We define a fuchsian affine action of a surface group to be such that the linear part factors through a representation of . We prove a fuchsian affine action of a surface group is never proper.
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones f…
We prove that any real-analytic, volume-preserving action of a lattice in a simple Lie group with $\Qrank(Γ)\geq 7$ on a closed 4-manifold of nonzero Euler characteristic factors through a finite group action.
By means of a Kaluza-Klein type argument we show that the Perelman's F-functional is the Einstein-Hilbert action in a space with extra ``phantom'' dimensions. In this way, we try to interpret some remarks of Perelman in the introduction and at the end of the first section in his first famous paper. As a consequence the…
Algorithm finds optimal regularizers for online linear optimization.
Study shows conditions for continuity of foliated homeomorphisms action on space of leaves.
We present an algorithm based on the \emph{Optimism in the Face of Uncertainty} (OFU) principle which is able to learn Reinforcement Learning (RL) modeled by Markov decision process (MDP) with finite state-action space efficiently. By evaluating the state-pair difference of the optimal bias function , the propos…
Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model. We propose a parametric Q-learning algorithm that finds an approximate-optimal policy using a sample size proportional to the feature dimension and invariant …
Let X be a simply connected compact Riemannian symmetric space, let U be the universal covering group of the identity component of the isometry group of X, and let \g denote the complexification of the Lie algebra of U, \g=\u^\C. Each \u-compatible triangular decomposition \g=\n_- + \h + \n_+ determines a Poisson Lie g…
We study the construction of quasimorphisms on groups acting on trees introduced by Monod and Shalom, that we call median quasimorphisms, and in particular we fully characterise actions on trees that give rise to non-trivial median quasimorphisms. Roughly speaking, either the action is highly transitive on geodesics, i…
We introduce the factored bandits model, which is a framework for learning with limited (bandit) feedback, where actions can be decomposed into a Cartesian product of atomic actions. Factored bandits incorporate rank-1 bandits as a special case, but significantly relax the assumptions on the form of the reward function…
Efficient Bayesian decision-making with intractable likelihoods.
Conditions for hyperbolic and relatively hyperbolic extensions of free groups using automorphisms with fixed points.
Improved exploration in factored average-reward MDPs reduces regret.
Any reinforcement learning algorithm that applies to all Markov decision processes (MDPs) will suffer regret on some MDP, where is the elapsed time and and are the cardinalities of the state and action spaces. This implies time to guarantee a near-optimal policy. In many settings…
We consider interactive learning and covering problems, in a setting where actions may incur different costs, depending on the response to the action. We propose a natural greedy algorithm for response-dependent costs. We bound the approximation factor of this greedy algorithm in active learning settings as well as in …