Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

164329493657 · Jun 202019922001200920172026
48 results for Discrete Action Space

Reinforcement learning (RL) in discrete action space is ubiquitous in real-world applications, but its complexity grows exponentially with the action-space dimension, making it challenging to apply existing on-policy gradient based deep RL algorithms efficiently. To effectively operate in multidimensional discrete acti…

2020-02-10abs ↗pdf ↗

Hybrid Policy Optimization tackles reinforcement learning in hybrid spaces, improving performance over PPO.

problem Credit assignment issues and biased gradients in hybrid discrete-continuous action spaces.
method Mixed gradient estimator combining pathwise and score-function gradients, reformulating problems in hybrid form.
result HPO substantially outperforms PPO on inventory control and switched systems, with performance gaps increasing with continuous action dimension.

Local-to-global principle for Morse actions on symmetric spaces.

problem Recognizing Morse actions on symmetric spaces.
method Equivariant Morse quasiisometric embeddings of trees into symmetric spaces.
result Algorithmic recognizability of Morse actions and construction of Morse Schottky subgroups.

Hybrid RL method optimizes trading by balancing continuous and discrete actions.

problem Optimal execution in algorithmic trading with continuous-discrete action space.
method Combines continuous and discrete RL agents for better trading decisions.
result Significantly outperforms existing methods in trading efficiency and stability.

It has long been assumed that high dimensional continuous control problems cannot be solved effectively by discretizing individual dimensions of the action space due to the exponentially large number of bins over which policies would have to be learned. In this paper, we draw inspiration from the recent success of sequ…

2017-05-14abs ↗pdf ↗

Being able to reason in an environment with a large number of discrete actions is essential to bringing reinforcement learning to a larger class of problems. Recommender systems, industrial plants and language models are only some of the many real-world tasks involving large numbers of discrete actions for which curren…

2015-12-24abs ↗pdf ↗

Investigates offline RL in factorisable action spaces, overcoming overestimation bias.

problem Overestimation bias in value estimates for unseen state-action pairs.
method Value-decomposition approach in DecQN, adapted for factorised discrete action spaces.
result Demonstrates the effectiveness of factorised approach in offline RL.

Let G/H be a strongly regular homogeneous space such that H is a Lie group of inner type. We show that G/H admits a proper action of a discrete non-virtually abelian subgroup of G if and only if G/H admits a proper action of a subgroup L of G locally isomorphic to SL(2,R). We classify all such spaces.

2015-01-28abs ↗pdf ↗

We study the geometry and dynamics of discrete infinite covolume subgroups of higher rank semisimple Lie groups. We introduce and prove the equivalence of several conditions, capturing "rank one behavior'' of discrete subgroups of higher rank Lie groups. They are direct generalizations of rank one equivalents to convex…

2014-03-29abs ↗pdf ↗

Paper solves POMDPs in continuous time and discrete spaces.

problem Optimal decision making in discrete state and action space systems under partial observability.
method Combining optimal filtering theory and deep learning to solve a Hamilton-Jacobi-Bellman equation.
result Derives a mathematical description and solution approach for continuous-time POMDPs.

This article gives an up-to-date account of the theory of discrete group actions on non-Riemannian homogeneous spaces. As an introduction of the motifs of this article, we begin by reviewing the current knowledge of possible global forms of pseudo-Riemannian manifolds with constant curvatures, and discuss what kind of …

2006-03-14abs ↗pdf ↗

Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an alternative version of the Soft Actor-Critic algorithm that is applicable to dis…

2019-10-16abs ↗pdf ↗

In this paper we propose a process of lagrangian reduction and reconstruction for nonholonomic discrete mechanical systems where the action of a continuous symmetry group makes the configuration space a principal bundle. The result of the reduction process is a discrete dynamical system that we call the discrete reduce…

2010-04-24abs ↗pdf ↗

Stochastic Q-learning tackles large action spaces with reduced computation.

problem Effective decision-making in complex environments with large discrete action spaces.
method Stochastic value-based RL approaches that consider a sublinear number of actions in each iteration.
result Stochastic Q-learning achieves near-optimal returns with significantly reduced computation time.

Generalizing a classical theorem of Carlson and Toledo, we prove that any Zariski dense isometric action of a Kähler group on the real hyperbolic space of dimension at least 3 factors through a homomorphism onto a cocompact discrete subgroup of PSL(2,R). We also study actions of Kähler groups on infinite dimensional re…

2010-12-07abs ↗pdf ↗

Study proves duality in exotic option pricing under uncertain model and delayed information.

problem Pricing and hedging of multi-action exotic options under nondominated model uncertainty and delayed information.
method Reformulated superhedging problem as a European option problem, proving duality results.
result Superhedging price equals model-based price with future look-up power.

A 2-manifold's group structure is deduced from orbit configuration spaces.

problem Understanding the fundamental groups of orbit configuration spaces.
method Relating the four-term exact sequence of orbifold pure braid groups to the fundamental groups of the orbit configuration spaces.
result Fundamental groups of orbit configuration spaces form a four-term exact sequence.

A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, na…

2018-05-24abs ↗pdf ↗

The theme of this survey is that subgroups of the mapping class group of a finite type surface S can be studied via the geometric/dynamical properties of their action on the Thurston compactification of the Teichmuller space of S, just as discrete subgroups of the isometries of hyperbolic space can be studied via their…

2007-02-14abs ↗pdf ↗

We study discrete, cocompact, isometric actions of groups on Hadamard spaces, and the induced actions on ideal boundaries. For a class of groups generalizing fundamental groups of three-dimensional graph manifolds, we find a set of invariants for the action which determine the boundary action up to equivariant homeomor…

1999-11-22abs ↗pdf ↗

We begin by showing that commensurators of Zariski dense subgroups of isometry groups of symmetric spaces of non-compact type are discrete provided that the limit set on the Furstenberg boundary is not invariant under the action of a (virtual) simple factor. In particular for rank one or simple Lie groups, Zariski dens…

2010-06-27abs ↗pdf ↗

In this short note, we prove that the space of all admissible piecewise linear metrics parameterized by length square on a triangulated manifolds is a convex cone. We further study Regge's Einstein-Hilbert action and give a much more reasonable definition of discrete Einstein metric than our former version in \cite{G}.…

2015-08-25abs ↗pdf ↗

ZoomRL learns efficient strategies for large state-action spaces using a metric.

problem Handling large state-action spaces in reinforcement learning.
method ZoomRL leverages continuous bandits to adaptively discretize the joint space.
result Achieves worst-case regret of $ ilde{O}(H^{ rac{5}{2}} K^{ rac{d+1}{d+2}})$.