Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

73146219292 · Jun 202019922001200920172026
48 results for representative actions

Simplifies large action space bandits by selecting representative actions.

problem Efficiently managing large action spaces with correlated outcomes.
method Random sampling and solving of bandit instances to identify representative actions.
result The algorithm selects a smaller set of representative actions that perform nearly as well as the full action space.

Solves action selection for large spaces in RL, achieving near-optimal performance.

problem Selecting a small, representative subset of actions from a large, shared action space.
method Extends meta-bandit approach to MDPs, using a relaxed sub-Gaussian process model.
result Achieves performance comparable to full action space, with theoretical guarantees.

Improved action recognition in live videos with hybrid FR-DL method.

problem High computational costs and lack of temporal information in conventional action recognition.
method Automated selection of representative frames, feature extraction, background subtraction, HOG, deep neural network, LSTM, Softmax-KNN classifier.
result Significant improvement in accuracy and speed compared to state-of-the-art methods.

CLIP dataset helps extract action items from hospital discharge notes.

problem Lost action items in long discharge notes hinder information sharing.
method Created CLIP dataset of annotated discharge notes, used multi-aspect extractive summarization, trained models on pre-trained language models and context.
result Best models improved by incorporating context and pre-trained language models.

We derive a consistent differential representation for the dynamics of a self-financing portfolio for different hedging strategies. In the basis of the derivation there is the so called "retarded action principle", which represents the causality in the evolution of dependent stochastic variables. We demonstrate this pr…

2015-09-30abs ↗pdf ↗

Orbifold groupoids have been recently widely used to represent both effective and ineffective orbifolds. We show that every orbifold groupoid can be faithfully represented on a continuous family of finite dimensional Hilbert spaces. As a consequence we obtain the result that every orbifold groupoid is Morita equivalent…

2007-09-03abs ↗pdf ↗

Approximate linear programming (ALP) represents one of the major algorithmic families to solve large-scale Markov decision processes (MDP). In this work, we study a primal-dual formulation of the ALP, and develop a scalable, model-free algorithm called bilinear ππ learning for reinforcement learning when a sampling or…

2018-04-27abs ↗pdf ↗

Improved Q-learning for multi-agent reinforcement learning by weighting joint action values.

problem QMIX restricts QQ-values to monotonic mixtures, limiting complex value functions.
method Introduced weighted projection to recover optimal policies, improving performance.
result CW QMIX and OW QMIX outperform baseline QMIX on multi-agent tasks.

Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…

2013-09-26abs ↗pdf ↗

We introduce a Hopf algebroid associated to a proper Lie group action on a smooth manifold. We prove that the cyclic cohomology of this Hopf algebroid is equal to the de Rham cohomology of invariant differential forms. When the action is cocompact, we develop a generalized Hodge theory for the de Rham cohomology of inv…

2010-02-23abs ↗pdf ↗

Intelligent agents can learn to represent the action spaces of other agents simply by observing them act. Such representations help agents quickly learn to predict the effects of their own actions on the environment and to plan complex action sequences. In this work, we address the problem of learning an agent's action…

2018-06-25abs ↗pdf ↗

Fine-grained action segmentation in long untrimmed videos is an important task for many applications such as surveillance, robotics, and human-computer interaction. To understand subtle and precise actions within a long time period, second-order information (e.g. feature covariance) or higher is reported to be effectiv…

2019-06-03abs ↗pdf ↗

Poisson and symplectic structures discussed in lecture notes.

problem Exploring Poisson and symplectic structures in mathematics.
method Presentation of Poisson and symplectic structures, group actions, moment maps, and phase space reduction.
result Comprehensive review of Poisson and symplectic structures, group actions, and reduction.

Diffusion-QL uses diffusion models to improve offline RL performance.

problem Offline RL struggles with function approximation errors on out-of-distribution actions.
method Diffusion-QL represents the policy as a conditional diffusion model and optimizes action-values.
result Diffusion-QL achieves state-of-the-art performance on D4RL benchmark tasks.

PFPN uses particle filtering to improve character control in physics-based simulations.

problem Premature commitment to suboptimal actions in high-dimensional continuous control problems for articulated characters.
method Proposes a particle-based action policy using particle filtering to dynamically explore and discretize the action space.
result Demonstrates better imitation performance and robustness to external perturbations compared to Gaussian policies.

SEMI uses multisensory incongruity to self-supervise exploration in reinforcement learning.

problem Efficient exploration in reinforcement learning with sparse or missing rewards.
method SEMI incentivizes exploration by maximizing multisensory incongruity, measured in perception and action incongruity.
result SEMI improves sample efficiency and learns skills without external rewards.

In 1974, Berezin proposed a quantum theory for dynamical systems having a Kähler manifold as their phase space. The system states were represented by holomorphic functions on the manifold. For any homogeneous Kähler manifold, the Lie algebra of its group of motions may be represented either by holomorphic differential …

1994-07-15abs ↗pdf ↗

A Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects. Early work in RMDPs outputs generalized (instance-independent) first-order policies or value functions as a means to solve all instanc…

2020-02-18abs ↗pdf ↗

Let GG be a finite group. Noncommutative geometry of unital GG-algebras is studied. A geometric structure is determined by a spectral triple on the crossed product algebra associated with the group action. This structure is to be viewed as a representative of a noncommutative orbifold. Based on a study of classical o…

2015-04-18abs ↗pdf ↗

Investigates sequential problems on graph structures and large action spaces.

problem Sequential decision-making on graph structures and large action spaces.
method Spectral bandits, side observations, influence maximization, kernel bandits, polymatroid bandits, function optimization, infinitely many-arms bandits.
result Contributions to graph and structured bandits.

Paper presents an action principle for Einstein-Weyl equations in 3D.

problem Finding an action principle for Einstein-Weyl equations.
method Metric affine f(R) gravity action plus additional terms involving Lagrange multipliers and gravitational Chern-Simons contributions.
result The Weyl vector dynamics is governed by a special case of the generalized monopole equation.

This is a review with examples concerning the concepts of affine (in particular, constant and linear) vector fields and fundamental vector fields on a manifold. The affine, linear and constant vector fields on a manifold are shown to be in a bijective correspondence with the fundamental vector fields on it of respectiv…

2006-02-01abs ↗pdf ↗

New RL method handles large state-action spaces with complex models.

problem Complex models and large state-action spaces in reinforcement learning.
method π-KRVI, an optimistic modification of least-squares value iteration using kernel ridge regression.
result First order-optimal regret guarantees under general settings, improving over state of the art.

New method recovers diverse policies from expert data using state-action pair weighting.

problem Recovering diverse policies from expert trajectories.
method Pointwise mutual information weighted behavioral cloning.
result Effective in focusing on state-action pairs most representative of the style.

PSI-LinUCB improves scalability for large recommender systems.

problem Efficiently training and inferring for large action spaces in recommender systems.
method Represent inverse design matrix as diagonal + low-rank correction, derive stable rank-1 and batched updates, use projector-splitting integrator.
result Demonstrated effectiveness on recommender system datasets, achieving scalable training and inference.

RANDPOL uses randomized networks for efficient reinforcement learning in continuous state and action MDPs.

problem Efficient reinforcement learning in environments with continuous state and action spaces.
method RANDPOL uses randomized function approximation to represent policy and value functions, providing finite time guarantees and improved numerical performance.
result RANDPOL achieves better numerical performance and provides finite time guarantees compared to deep neural network based algorithms.

Graphon game model simplifies stochastic interactions among agents.

problem Complex interactions among heterogeneous agents in stochastic games.
method Introduced a discrete-time graphon game formulation with a representative player.
result Existence and uniqueness of graphon equilibrium proven with mild assumptions.

A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a de…

2020-02-05abs ↗pdf ↗

TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.

problem Stochastic linear bandits with partially observed actions in settings like recommendation and healthcare.
method TOFU-POV estimates latent action subspace, imputes missing actions, and runs OFUL in low-dimensional coordinates.
result TOFU-POV achieves T\sqrt{T} regret scaling with intrinsic subspace dimension, improving upon natural baselines.

Using the twistor correspondence, this article gives a one-to-one correspondence between germs of toric anti-self-dual conformal classes and certain holomorphic data determined by the induced action on twistor space. Recovering the metric from the holomorphic data leads to the classical problem of prescribing the Cech …

2006-02-20abs ↗pdf ↗