A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
DRL agents perform poorly at high decision frequencies, but a new algorithm improves performance.
problem DRL agents struggle at high decision frequencies, leading to poor performance.
method Proved that DRL agents' action-conditioned return distributions collapse to their policy's return distribution as decision frequency increases. Defined superiority as a probabilistic generalization of advantage for high-frequency value-based RL.
result Proper modeling of superiority distribution improves performance of controllers at high decision frequencies.
In an effort to better understand the different ways in which the discount factor affects the optimization process in reinforcement learning, we designed a set of experiments to study each effect in isolation. Our analysis reveals that the common perception that poor performance of low discount factors is caused by (to…
We survey the use of dynamics of SL(2,R)-actions to understand gap distributions for various sequences of subsets of [0,1), particularly those arising from special trajectories of various two-dimensional dynamical systems. We state and prove an abstract theorem that gives a unified explanation for some of the ex…
Unpublished results of S Straus and W Browder state that two notions of homotopy equivalence for manifolds with smooth group actions - isovariant and equivariant - often coincide under a condition called the Gap Hypothesis; the proofs use deep results in geometric topology. This paper analyzes the difference between th…
The harmonic action functional allows a natural generalisation to semi-Riemannian supergeometry, referred to as superharmonic action, which resembles the supersymmetric sigma models studied in high energy physics. We show that Killing vector fields are infinitesimal supersymmetries of the superharmonic action and prove…
We prove that for any isometric action of a group on a unit sphere of dimension larger than one, the quotient space has diameter zero or larger than a universal dimension-independent positive constant.
Skeletal signatures were introduced in [J W Anderson and A Wootton, A Lower Bound for the Number of Group Actions on a Compact Riemann Surface, Algebr. Geom. Topol. 12 (2012) 19--35.] as a tool to describe the space of all signatures with which a group can act on a surface of genus σ≥2. In the present paper we pr…
Given a constant magnetic field on Euclidean space Rp determined by a skew-symmetric (p×p) matrix Θ, and a Zp-invariant probability measure μ on the disorder set Σ which is by hypothesis a Cantor set, where the action is assumed to be minimal, the corresponding Integrated Density…
Life-expectancy is a complex outcome driven by genetic, socio-demographic, environmental and geographic factors. Increasing socio-economic and health disparities in the United States are propagating the longevity-gap, making it a cause for concern. Earlier studies have probed individual factors but an integrated pictur…
Let Γ′<Γ be two discrete groups acting properly by isometries on a Gromov-hyperbolic space X. We prove that their critical exponents coincide if and only if Γ′ is co-amenable in Γ, under the assumption that the action of Γ on X is strongly positively recurrent, i.e. has a growth gap at infinity. This genera…
A gap in the proof of the main result in reference [1] in our original submission propagated into the constructions presented in the first version of our manuscript. In this version we give an alternative proof for the existence of Riemannian metrics with positive Ricci curvature on an infinite subfamily of closed, sim…
We study the dynamics of the Teichmuller flow in the moduli space of Abelian differentials (and more generally, its restriction to any connected component of a stratum). We show that the (Masur-Veech) absolutely continuous invariant probability measure is exponentially mixing for the class of Holder observables. A geom…
Deep Reinforcement Learning (DRL) has been applied to address a variety of cooperative multi-agent problems with either discrete action spaces or continuous action spaces. However, to the best of our knowledge, no previous work has ever succeeded in applying DRL to multi-agent problems with discrete-continuous hybrid (…
We propose a new low-cost machine-learning-based methodology which assists designers in reducing the gap between the problem and the solution in the design process. Our work applies reinforcement learning (RL) to find the optimal task-oriented design solution through the construction of the design action for each task.…
This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervision, LfO is more practical in leveraging previously inapplicable resources (e.g. videos), yet more challenging…
We address reinforcement learning problems with finite state and action spaces where the underlying MDP has some known structure that could be potentially exploited to minimize the exploration rates of suboptimal (state, action) pairs. For any arbitrary structure, we derive problem-specific regret lower bounds satisfie…
New approach makes deep reinforcement learning robust without assuming adversary knowledge.
problem Deep reinforcement learning policies are vulnerable to state observation perturbations.
method Proposes an adversary agnostic robust DRL paradigm using policy distillation with two terms: prescription gap maximization and Jacobian regularization.
result Boosts adversarial robustness on five Atari games compared to state-of-the-art methods.
We consider a new family of operators for reinforcement learning with the goal of alleviating the negative effects and becoming more robust to approximation or estimation errors. Various theoretical results are established, which include showing on a sample path basis that our family of operators preserve optimality an…
We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…