Simplifies large action space bandits by selecting representative actions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper presents a preliminary study comparing different observation and action space representations for Deep Reinforcement Learning (DRL) in the context of Real-time Strategy (RTS) games. Specifically, we compare two representations: (1) a global representation where the observation represents the whole game state…
Solves action selection for large spaces in RL, achieving near-optimal performance.
We propose a framework for modeling and estimating the state of controlled dynamical systems, where an agent can affect the system through actions and receives partial observations. Based on this framework, we propose the Predictive State Representation with Random Fourier Features (RFFPSR). A key property in RFF-PSRs …
Improved action recognition in live videos with hybrid FR-DL method.
New tools prove smooth actions on exotic spheres.
In this paper we propose the use of quantum genetic algorithm to optimize the support vector machine (SVM) for human action recognition. The Microsoft Kinect sensor can be used for skeleton tracking, which provides the joints' position data. However, how to extract the motion features for representing the dynamics of a…
CLIP dataset helps extract action items from hospital discharge notes.
Classifies orbits of Hurwitz actions on dihedral quandles.
Recognizing group activities is challenging due to the difficulties in isolating individual entities, finding the respective roles played by the individuals and representing the complex interactions among the participants. Individual actions and group activities in videos can be represented in a common framework as the…
We derive a consistent differential representation for the dynamics of a self-financing portfolio for different hedging strategies. In the basis of the derivation there is the so called "retarded action principle", which represents the causality in the evolution of dependent stochastic variables. We demonstrate this pr…
Orbifold groupoids have been recently widely used to represent both effective and ineffective orbifolds. We show that every orbifold groupoid can be faithfully represented on a continuous family of finite dimensional Hilbert spaces. As a consequence we obtain the result that every orbifold groupoid is Morita equivalent…
Approximate linear programming (ALP) represents one of the major algorithmic families to solve large-scale Markov decision processes (MDP). In this work, we study a primal-dual formulation of the ALP, and develop a scalable, model-free algorithm called bilinear learning for reinforcement learning when a sampling or…
Improved Q-learning for multi-agent reinforcement learning by weighting joint action values.
Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…
New insights into the geometry of flows on 3-manifolds.
Proves conjecture simplifying mapping class group action on Steinberg module.
The optimal policy of a reinforcement learning problem is often discontinuous and non-smooth. I.e., for two states with similar representations, their optimal policies can be significantly different. In this case, representing the entire policy with a function approximator (FA) with shared parameters for all states may…
New method constructs actions on R capturing foliations, proving left-orderability.
We introduce a Hopf algebroid associated to a proper Lie group action on a smooth manifold. We prove that the cyclic cohomology of this Hopf algebroid is equal to the de Rham cohomology of invariant differential forms. When the action is cocompact, we develop a generalized Hodge theory for the de Rham cohomology of inv…
Intelligent agents can learn to represent the action spaces of other agents simply by observing them act. Such representations help agents quickly learn to predict the effects of their own actions on the environment and to plan complex action sequences. In this work, we address the problem of learning an agent's action…
A new method learns action representations for reinforcement learning.
Compact groups can be represented as dessin automorphisms.
Fine-grained action segmentation in long untrimmed videos is an important task for many applications such as surveillance, robotics, and human-computer interaction. To understand subtle and precise actions within a long time period, second-order information (e.g. feature covariance) or higher is reported to be effectiv…
Poisson and symplectic structures discussed in lecture notes.
Representation of human actions as a sequence of human body movements or action attributes enables the development of models for human activity recognition and summarization. We present an extension of the low-rank representation (LRR) model, termed the clustering-aware structure-constrained low-rank representation (CS…
Diffusion-QL uses diffusion models to improve offline RL performance.
PFPN uses particle filtering to improve character control in physics-based simulations.
SEMI uses multisensory incongruity to self-supervise exploration in reinforcement learning.
In 1974, Berezin proposed a quantum theory for dynamical systems having a Kähler manifold as their phase space. The system states were represented by holomorphic functions on the manifold. For any homogeneous Kähler manifold, the Lie algebra of its group of motions may be represented either by holomorphic differential …
A Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects. Early work in RMDPs outputs generalized (instance-independent) first-order policies or value functions as a means to solve all instanc…
Let be a closed 4-manifold with a free circle action. If the orbit manifold satisfies an appropriate fibering condition, then we show how to represent a cone in by symplectic forms. This generalizes earlier constructions by Thurston, Bouyakoub and Fernández-Gray-Morgan. In the case that is the…
Let be a finite group. Noncommutative geometry of unital -algebras is studied. A geometric structure is determined by a spectral triple on the crossed product algebra associated with the group action. This structure is to be viewed as a representative of a noncommutative orbifold. Based on a study of classical o…
Investigates sequential problems on graph structures and large action spaces.
Paper presents an action principle for Einstein-Weyl equations in 3D.
This is a review with examples concerning the concepts of affine (in particular, constant and linear) vector fields and fundamental vector fields on a manifold. The affine, linear and constant vector fields on a manifold are shown to be in a bijective correspondence with the fundamental vector fields on it of respectiv…
Portfolio traders strive to identify dynamic portfolio allocation schemes so that their total budgets are efficiently allocated through the investment horizon. This study proposes a novel portfolio trading strategy in which an intelligent agent is trained to identify an optimal trading action by using deep Q-learning. …
New RL method handles large state-action spaces with complex models.
As Computer Vision moves from a passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will …
New framework for conformal equivariant cycles in KK-theory.
New method recovers diverse policies from expert data using state-action pair weighting.
PSI-LinUCB improves scalability for large recommender systems.
RANDPOL uses randomized networks for efficient reinforcement learning in continuous state and action MDPs.
Graphon game model simplifies stochastic interactions among agents.
A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a de…
Extends distribution algebra concept to Lie groupoids.
TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.
Using the twistor correspondence, this article gives a one-to-one correspondence between germs of toric anti-self-dual conformal classes and certain holomorphic data determined by the induced action on twistor space. Recovering the metric from the holomorphic data leads to the classical problem of prescribing the Cech …