A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
In this paper, we study multi-armed bandit problems in explore-then-commit setting. In our proposed explore-then-commit setting, the goal is to identify the best arm after a pure experimentation (exploration) phase and exploit it once or for a given finite number of times. We identify that although the arm with the hig…
In classical reinforcement learning, when exploring an environment, agents accept arbitrary short term loss for long term gain. This is infeasible for safety critical applications, such as robotics, where even a single unsafe action may cause system failure. In this paper, we address the problem of safely exploring fin…
This paper studies a recent proposal to use randomized value functions to drive exploration in reinforcement learning. These randomized value functions are generated by injecting random noise into the training data, making the approach compatible with many popular methods for estimating parameterized value functions. B…
We study the exploration problem in episodic MDPs with rich observations generated from a small number of latent states. Under certain identifiability assumptions, we demonstrate how to estimate a mapping from the observations to latent states inductively through a sequence of regression and clustering steps -- where p…
We present a 1-parameter family of finite action solutions to the S0(2,1) Hitchin's equations and explore some of its basic properties. For a fixed value of the parameter, the solution is smooth. We conclude by showing a multi-particle generalization of our basic solutions.
We examine groups whose resonance varieties, characteristic varieties and Sigma-invariants have a natural arithmetic group symmetry, and we explore implications on various finiteness properties of subgroups. We compute resonance varieties, characteristic varieties and Alexander polynomials of Torelli groups, and we sho…
We consider the class of curves of finite total curvature, as introduced by Milnor. This is a natural class for variational problems and geometric knot theory, and since it includes both smooth and polygonal curves, its study shows us connections between discrete and differential geometry. To explore these ideas, we co…
This paper approaches the definition and properties of dynamic convex risk measures through the notion of a family of concave valuation operators satisfying certain simple and credible axioms. Exploring these in the simplest context of a finite time set and finite sample space, we find natural risk-transfer and time-co…
Contextual bandits have the same exploration-exploitation trade-off as standard multi-armed bandits. On adding positive externalities that decay with time, this problem becomes much more difficult as wrong decisions at the start are hard to recover from. We explore existing policies in this setting and highlight their …
In this paper, we explore holomorphic Segre preserving maps. First, we investigate holomorphic Segre preserving maps sending the complexification M of a generic real analytic submanifold $M \subseteq \C^N$ of finite type at some point p into the complexification M′ of a generic real analytic s…
A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-learning with…
Control of non-episodic, finite-horizon dynamical systems with uncertain dynamics poses a tough and elementary case of the exploration-exploitation trade-off. Bayesian reinforcement learning, reasoning about the effect of actions and future observations, offers a principled solution, but is intractable. We review, then…
This paper studies the problem of identifying any k distinct arms among the top ρ fraction (e.g., top 5\%) of arms from a finite or infinite set with a probably approximately correct (PAC) tolerance ε. We consider two cases: (i) when the threshold of the top arms' expected rewards is known and (ii) when it is unk…