Survey of three geometric frameworks for action-dependent field theories.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Policy gradient methods have enjoyed great success in deep reinforcement learning but suffer from high variance of gradient estimates. The high variance problem is particularly exasperated in problems with long horizons or high-dimensional action spaces. To mitigate this issue, we derive a bias-free action-dependent ba…
Black-box optimizers that explore in parameter space have often been shown to outperform more sophisticated action space exploration methods developed specifically for the reinforcement learning problem. We examine these black-box methods closely to identify situations in which they are worse than action space explorat…
Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces variance and improv…
We consider a sequential learning problem with Gaussian payoffs and side information: after selecting an action , the learner receives information about the payoff of every action in the form of Gaussian observations whose mean is the same as the mean payoff, but the variance depends on the pair (and may…
Paper proposes a new HMM approach for better action recognition.
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can become arbitrarily large as the number of actions increases. We establish new bounds that depend instead on a notion of rate-distortion. Among o…
We consider interactive learning and covering problems, in a setting where actions may incur different costs, depending on the response to the action. We propose a natural greedy algorithm for response-dependent costs. We bound the approximation factor of this greedy algorithm in active learning settings as well as in …
We give a diameter bound for fundamental domains for isometric actions of the fundamental group of a closed hyperbolic surface on a delta-hyperbolic space, where the bound depends on the hyperbolicity constant delta, the genus of the surface, and the injectivity radius of the action, which we assume to be strictly posi…
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
We exhibit large classes of local actions for the vacuum Einstein equations. In presence of fermions, or more generally of matter which couple to the connection, these actions lead to inequivalent equations revealing an arbitrary number of parameters. Even in the pure gravitational sector, any corresponding quantum the…
We discuss how the global geometry and topology of manifolds depend on different group actions of their fundamental groups, and in particular, how properties of a non-trivial compact 4-dimensional cobordism whose interior has a complete hyperbolic structure depend on properties of the variety of discrete representa…
PQR estimates reward functions from actions and states without assuming state-only rewards.
Proposes a neural network for efficient deep hedging strategies.
Develops a new reinforcement learning framework for complex control problems.
New algorithm reduces regret in private online learning with optimal gap-dependent rate.
User engagement in social networks depends critically on the number of online actions their users take in the network. Can we design an algorithm that finds when to incentivize users to take actions to maximize the overall activity in a social network? In this paper, we model the number of online actions over time usin…
Survey on finite group actions on manifolds.
New algorithm for bandits with delayed action effects, reducing regret.
This paper provides theoretical foundations for using quantized actions in behavior cloning.
Develops a regression approach for solving MDPs with general state and action spaces.
Unified framework for corruption-robust linear bandits with optimal gap-dependent misspecification bounds.
We introduce a rich class of graphical models for multi-armed bandit problems that permit both the state or context space and the action space to be very large, yet succinctly specify the payoffs for any context-action pair. Our main result is an algorithm for such models whose regret is bounded by the number of parame…
New algorithm tackles multi-agent reinforcement learning with optimal convergence rate.
New algorithm tackles nonstationary linear bandits with latent dynamics.
New algorithm identifies optimal actions in large reward spaces efficiently.
We present and study a partial-information model of online learning, where a decision maker repeatedly chooses from a finite set of actions, and observes some subset of the associated losses. This naturally models several situations where the losses of different actions are related, and knowing the loss of one action p…
Study of symplectomorphisms on ruled surfaces under circle actions.
We derive a consistent differential representation for the dynamics of a self-financing portfolio for different hedging strategies. In the basis of the derivation there is the so called "retarded action principle", which represents the causality in the evolution of dependent stochastic variables. We demonstrate this pr…
Functional determinant for mixed signature sphere products depends on sphere dimensions and parity.
We show that J. Lott's equivariant higher analytic torsion for compact group actions depends only on the equivariant Euler characteristic.
Describes reconstructing Poisson structures from Lie group actions.
Algorithm maximizes rewards with a budget and giving up option.
New method for linear bandits with unknown sparsity, improving sparse regret bounds.
Finite p-group actions on manifolds have limited stabilizer subgroups.
We develop a normative framework for hierarchical model-based policy optimization based on applying second-order methods in the space of all possible state-action paths. The resulting natural path gradient performs policy updates in a manner which is sensitive to the long-range correlational structure of the induced st…
We consider the reduced Allen-Cahn action functional, which appears as the sharp interface limit of the Allen-Cahn action functional and can be understood as a formal action functional for a stochastically perturbed mean curvature flow. For suitable evolutions of generalized hypersurfaces this functional consists of th…
Proposes MDR estimator for unbiased OPE with large action spaces.
Rigidity of elliptic genera proven for non-spin manifolds with -action.
Extends integrability to cosymplectic manifolds.
We compute the quotient of the self-duality equation for conformal metrics by the action of the diffeomorphism group. We also determine Hilbert polynomial, counting the number of independent scalar differential invariants depending on the jet-order, and the corresponding Poincaré function. We describe the field of rati…
Let G be a group and let M be a CAT(0) proper metric space (e.g. a simply connected complete Riemannian manifold of non-positive sectional curvature or a locally finite tree). Isometric actions of G on M are (by definition) points in the space R := Hom(G, Isom(M)) with the compact open topology. Sample theorems: 1. The…
Improved reinforcement learning for episodes with varying action sets.
SPEDER extracts state-action abstraction from dynamics for reinforcement learning.
The paper calculates variations of Einstein-Hilbert action on CR manifolds.
This is the second of two papers but has been written so as to have minimal dependence on the first paper (which is also on this archive). Let G be a group and let M be a CAT(0) proper metric space (e.g. a simply connected complete Riemannian manifold of non-positive sectional curvature or a locally finite tree). Assum…
Complex activity recognition is challenging due to the inherent uncertainty and diversity of performing a complex activity. Normally, each instance of a complex activity has its own configuration of atomic actions and their temporal dependencies. We propose in this paper an atomic action-based Bayesian model that const…
Let be a semisimple Lie group with all simple factors of real rank at least two. Let be a lattice. We prove a very general local rigidity result about actions of or . This shows that almost all so-called "standard actions" are locally rigid. As a special case, we see that any action of by toral aut…