A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent has to explore unknown actions, some of which can be bad, to learn better action…
We discuss multi-task online learning when a decision maker has to deal simultaneously with M tasks. The tasks are related, which is modeled by imposing that the M-tuple of actions taken by the decision maker needs to satisfy certain constraints. We give natural examples of such restrictions and then discuss a general …
In standard reinforcement learning (RL), a learning agent seeks to optimize the overall reward. However, many key aspects of a desired behavior are more naturally expressed as constraints. For instance, the designer may want to limit the use of unsafe actions, increase the diversity of trajectories to enable exploratio…
S.Bauer and M.Furuta defined a stable cohomotopy refinement of the Seiberg-Witten invariants. In this paper, we prove a vanishing theorem of Bauer-Furuta invariants for 4-manifolds with smooth Z/2-actions. As an application, we give a constraint on smooth Z/2-actions on homotopy K3#K3, and construct a nonsmoothable loc…
In classical reinforcement learning, when exploring an environment, agents accept arbitrary short term loss for long term gain. This is infeasible for safety critical applications, such as robotics, where even a single unsafe action may cause system failure. In this paper, we address the problem of safely exploring fin…
Consider an effective Hamiltonian torus action T×M→M on a topologically twisted,generalized complex manifold M of dimension 2n. We prove that the rank(T)≤n−2 and that the topological twisting survives Hamiltonian reduction. We then construct a large new class of such actions satisfying $rank(T) =…
We study the effect of impairment on stochastic multi-armed bandits and develop new ways to mitigate it. Impairment effect is the phenomena where an agent only accrues reward for an action if they have played it at least a few times in the recent past. It is practically motivated by repetition and recency effects in do…
Neural Index Policy for multi-action bandits with heterogeneous budgets.
problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.
Autonomous cyber-physical agents and systems play an increasingly large role in our lives. To ensure that agents behave in ways aligned with the values of the societies in which they operate, we must develop techniques that allow these agents to not only maximize their reward in an environment, but also to learn and fo…
In this paper, we present a new task that investigates how people interact with and make judgments about towers of blocks. In Experiment~1, participants in the lab solved a series of problems in which they had to re-configure three blocks from an initial to a final configuration. We recorded whether they used one hand …
For a Lie group G and a smooth manifold W, we study the difference between smooth actions of G on W and bundles over the classifying space of G with fiber W and structure group Diff(W). In particular, we exhibit smooth manifold bundles over BSU(2) that are not induced by an action. The main tool for reaching this goal …
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones f…
Using the theory of group action, we first introduce the concept of the automorphism group of an exponential family or a graphical model, thus formalizing the general notion of symmetry of a probabilistic model. This automorphism group provides a precise mathematical framework for lifted inference in the general expone…
We study continuous action reinforcement learning problems in which it is crucial that the agent interacts with the environment only through safe policies, i.e.,~policies that do not take the agent to undesirable situations. We formulate these problems as constrained Markov decision processes (CMDPs) and present safe p…
When the vacuum Einstein equations are cast in the form of hamiltonian evolution equations, the initial data lie in the cotangent bundle of the manifold MΣ of riemannian metrics on a Cauchy hypersurface Σ. As in every lagrangian field theory with symmetries, the initial data must satisfy constraints. But, unlike those …
We analyse the definition of quasi-local energy in GR based on a Hamiltonian analysis of the Einstein-Hilbert action initiated by Brown-York. The role of the constraint equations, in particular the Hamiltonian constraint on the timelike boundary, neglected in previous studies, is emphasized here. We argue that a consis…
This paper examines a proposal for gauging non-linear sigma models with respect to a Lie algebroid action. The general conditions for gauging a non-linear sigma model with a set of involutive vector fields are given. We show that it is always possible to find a set of vector fields which will (locally) admit a Lie alge…