Algorithm minimizes loss and constraint violations in online convex optimization with smooth penalties.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes Constrained Q-learning for reinforcement learning with constraints.
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent has to explore unknown actions, some of which can be bad, to learn better action…
Paper proposes a deep RL algorithm that can learn from forbidden actions.
New RL algorithm learns good actions from offline data, reducing uncertainty and divergence.
Two new algorithms optimize rewards while respecting safety constraints in sequential decisions.
Paper proposes Vertex Networks for reinforcement learning of control systems with safety guarantees.
Algorithm reduces regret in safe Bayesian optimization with monotonicity constraints.
Algorithm minimizes regret while adhering to unknown safety constraints.
Hybrid SAC improves RL for video games with discrete, continuous actions.
We discuss multi-task online learning when a decision maker has to deal simultaneously with M tasks. The tasks are related, which is modeled by imposing that the M-tuple of actions taken by the decision maker needs to satisfy certain constraints. We give natural examples of such restrictions and then discuss a general …
New algorithms optimize actions under time-varying constraints without projecting.
In standard reinforcement learning (RL), a learning agent seeks to optimize the overall reward. However, many key aspects of a desired behavior are more naturally expressed as constraints. For instance, the designer may want to limit the use of unsafe actions, increase the diversity of trajectories to enable exploratio…
S.Bauer and M.Furuta defined a stable cohomotopy refinement of the Seiberg-Witten invariants. In this paper, we prove a vanishing theorem of Bauer-Furuta invariants for 4-manifolds with smooth Z/2-actions. As an application, we give a constraint on smooth Z/2-actions on homotopy K3#K3, and construct a nonsmoothable loc…
New algorithm reduces regret and constraint violation in constrained bandit problems.
Safe reinforcement learning tackles safety constraints with linear approximations.
The paper introduces an adjacency constraint to improve goal-conditioned HRL.
Safe Gaussian Process Bandit Optimization with sub-linear regret bounds.
In classical reinforcement learning, when exploring an environment, agents accept arbitrary short term loss for long term gain. This is infeasible for safety critical applications, such as robotics, where even a single unsafe action may cause system failure. In this paper, we address the problem of safely exploring fin…
Consider an effective Hamiltonian torus action on a topologically twisted,generalized complex manifold of dimension . We prove that the and that the topological twisting survives Hamiltonian reduction. We then construct a large new class of such actions satisfying $rank(T) =…
New method optimizes portfolios by dynamically integrating ESG constraints.
Study circle actions on manifolds with 3 fixed points, finding dimension constraints and unique structures.
Proposes deep optimal feedback control for continuous-time systems with action constraints.
Bayesian inference over admissible histories leads to irreversible kinetics.
We study the effect of impairment on stochastic multi-armed bandits and develop new ways to mitigate it. Impairment effect is the phenomena where an agent only accrues reward for an action if they have played it at least a few times in the recent past. It is practically motivated by repetition and recency effects in do…
Study introduces constrained neural units over time for learning problems.
A new method for CMDP solving without compromising safety constraints.
Safe linear bandits over unknown polytopes avoid safety violations and suboptimal actions.
Suspensions of manifolds by circle surgeries are key in free action constructions.
Paper improves COCO problem, reducing constraint violation at the cost of slightly more regret.
New algorithm reduces constraint violation to while maintaining regret.
Neural Index Policy for multi-action bandits with heterogeneous budgets.
Autonomous cyber-physical agents and systems play an increasingly large role in our lives. To ensure that agents behave in ways aligned with the values of the societies in which they operate, we must develop techniques that allow these agents to not only maximize their reward in an environment, but also to learn and fo…
In this paper, we present a new task that investigates how people interact with and make judgments about towers of blocks. In Experiment~1, participants in the lab solved a series of problems in which they had to re-configure three blocks from an initial to a final configuration. We recorded whether they used one hand …
For a Lie group G and a smooth manifold W, we study the difference between smooth actions of G on W and bundles over the classifying space of G with fiber W and structure group Diff(W). In particular, we exhibit smooth manifold bundles over BSU(2) that are not induced by an action. The main tool for reaching this goal …
Robustness of Deep Reinforcement Learning (DRL) algorithms towards adversarial attacks in real world applications such as those deployed in cyber-physical systems (CPS) are of increasing concern. Numerous studies have investigated the mechanisms of attacks on the RL agent's state space. Nonetheless, attacks on the RL a…
Algorithm reduces episode count for CMDPs with constraints.
Optimal algorithm for submodular maximization in distributed networked scenarios.
Algorithm designs neural group actions for symmetric transformations.
We explore value-based solutions for multi-agent reinforcement learning (MARL) tasks in the centralized training with decentralized execution (CTDE) regime popularized recently. However, VDN and QMIX are representative examples that use the idea of factorization of the joint action-value function into individual ones f…
Using the theory of group action, we first introduce the concept of the automorphism group of an exponential family or a graphical model, thus formalizing the general notion of symmetry of a probabilistic model. This automorphism group provides a precise mathematical framework for lifted inference in the general expone…
We study continuous action reinforcement learning problems in which it is crucial that the agent interacts with the environment only through safe policies, i.e.,~policies that do not take the agent to undesirable situations. We formulate these problems as constrained Markov decision processes (CMDPs) and present safe p…
When the vacuum Einstein equations are cast in the form of hamiltonian evolution equations, the initial data lie in the cotangent bundle of the manifold MΣ of riemannian metrics on a Cauchy hypersurface Σ. As in every lagrangian field theory with symmetries, the initial data must satisfy constraints. But, unlike those …
The paper classifies constraint mappings in optimization problems.
New framework guides resource usage to achieve sublinear regret in adversarial settings.
We analyse the definition of quasi-local energy in GR based on a Hamiltonian analysis of the Einstein-Hilbert action initiated by Brown-York. The role of the constraint equations, in particular the Hamiltonian constraint on the timelike boundary, neglected in previous studies, is emphasized here. We argue that a consis…
This paper examines a proposal for gauging non-linear sigma models with respect to a Lie algebroid action. The general conditions for gauging a non-linear sigma model with a set of involutive vector fields are given. We show that it is always possible to find a set of vector fields which will (locally) admit a Lie alge…