Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,786 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Sep 199319922001200920172026
48 results for action constraints

Algorithm minimizes loss and constraint violations in online convex optimization with smooth penalties.

problem Minimizing loss and constraint violations in online convex optimization with smooth penalties.
method Projected gradient descent over a set around the current action.
result Both dynamic regret and constraint violation are bounded by the path-length.

Proposes Constrained Q-learning for reinforcement learning with constraints.

problem Optimizing multiple objectives while adhering to constraints in reinforcement learning.
method Directly restricts the action space in Q-update to learn optimal Q-function for constrained MDP.
result Improves safety and optimality in high-level decision making for autonomous driving.

Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…

2019-08-16abs ↗pdf ↗

In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent has to explore unknown actions, some of which can be bad, to learn better action…

2018-06-03abs ↗pdf ↗

New RL algorithm learns good actions from offline data, reducing uncertainty and divergence.

problem Limited applicability of current RL algorithms in real-world settings due to high costs of exploration.
method Proposes an algorithm for batch RL using a fixed offline dataset, with penalties for policy and value constraints.
result Compared favorably to state-of-the-art methods on 32 continuous-action benchmarks.

Two new algorithms optimize rewards while respecting safety constraints in sequential decisions.

problem Optimizing rewards with safety constraints in sequential decisions.
method Stage-wise conservative linear Thompson Sampling (SCLTS) and stage-wise conservative linear UCB (SCLUCB).
result Probabilistic regret bounds of order O(\sqrt{T} \log^{3/2}T) and O(\sqrt{T} \log T).

Paper proposes Vertex Networks for reinforcement learning of control systems with safety guarantees.

problem Challenges in reinforcement learning with hard state and action constraints.
method Vertex Networks incorporate safety constraints into policy network architecture, ensuring safety during exploration.
result Proposed Vertex Networks outperform vanilla reinforcement learning in benchmark control tasks.

Algorithm reduces regret in safe Bayesian optimization with monotonicity constraints.

problem Sequentially maximize unknown function with safety constraints.
method Sequential algorithms using Gaussian processes with safety constraints modeled as monotonicity.
result Sublinear regret achieved for expanding safe region and finding optimal ss.

Hybrid SAC improves RL for video games with discrete, continuous actions.

problem Improving RL performance in video games with practical constraints.
method Extension of Soft Actor-Critic (SAC) for handling discrete, continuous, and parameterized actions.
result Hybrid SAC successfully solves a high-speed driving task and is competitive on parameterized actions benchmarks.

We discuss multi-task online learning when a decision maker has to deal simultaneously with M tasks. The tasks are related, which is modeled by imposing that the M-tuple of actions taken by the decision maker needs to satisfy certain constraints. We give natural examples of such restrictions and then discuss a general …

2009-02-20abs ↗pdf ↗

New algorithms optimize actions under time-varying constraints without projecting.

problem Optimizing actions under time-varying constraints without projecting.
method Projection-free algorithms using linear optimization oracle.
result Guaranteed ildeO(T3/4) ilde{O}(T^{3/4}) regret and O(T7/8)O(T^{7/8}) constraints violation.

In standard reinforcement learning (RL), a learning agent seeks to optimize the overall reward. However, many key aspects of a desired behavior are more naturally expressed as constraints. For instance, the designer may want to limit the use of unsafe actions, increase the diversity of trajectories to enable exploratio…

2019-06-21abs ↗pdf ↗

S.Bauer and M.Furuta defined a stable cohomotopy refinement of the Seiberg-Witten invariants. In this paper, we prove a vanishing theorem of Bauer-Furuta invariants for 4-manifolds with smooth Z/2-actions. As an application, we give a constraint on smooth Z/2-actions on homotopy K3#K3, and construct a nonsmoothable loc…

2007-05-11abs ↗pdf ↗

New algorithm reduces regret and constraint violation in constrained bandit problems.

problem Optimizing under budget and stochastic constraints in resource-constrained settings.
method Lyapunov optimization methodology, tLyOn{ t LyOn} algorithm.
result Achieves O(KBlogB)O(\sqrt{K B\log B}) regret and zero constraint-violation for large BB.

Safe reinforcement learning tackles safety constraints with linear approximations.

problem Ensuring safety in reinforcement learning without violating constraints.
method Modeling safety as a linear cost function, developing SLUCB-QVI and RSLUCB-QVI algorithms for MDPs with linear function approximation.
result Achieved a nearly optimal regret bound for safe reinforcement learning, matching state-of-the-art unsafe algorithms.

The paper introduces an adjacency constraint to improve goal-conditioned HRL.

problem Training inefficiency in goal-conditioned HRL due to large action space.
method Restricting the high-level action space to a k-step adjacent region of the current state.
result The adjacency constraint preserves optimal hierarchical policies and improves HRL performance.

Consider an effective Hamiltonian torus action T×MMT\times M \to M on a topologically twisted,generalized complex manifold MM of dimension 2n2n. We prove that the rank(T)n2rank(T) \leq n-2 and that the topological twisting survives Hamiltonian reduction. We then construct a large new class of such actions satisfying $rank(T) =…

2009-04-07abs ↗pdf ↗

New method optimizes portfolios by dynamically integrating ESG constraints.

problem Static ESG scores mismatch sequential portfolio decisions.
method MACF-X, a family of adapters that learns ESG costs from multimodal evidence.
result Reduces tail ESG budget pressure while maintaining financial performance.

Study circle actions on manifolds with 3 fixed points, finding dimension constraints and unique structures.

problem Characterize circle actions on oriented manifolds with exactly 3 fixed points.
method Analyzes manifold dimensions, isotropy submanifolds, and uses quaternionic projective space as a reference.
result For a manifold with three fixed points, its dimension must be a multiple of 4, and specific weights are unique.

Proposes deep optimal feedback control for continuous-time systems with action constraints.

problem Learning optimal feedback control laws for robotic applications.
method Exploits Hamilton-Jacobi-Bellman equation and deep differential networks to learn optimal value function and feedback policy.
result Enables learning an optimal feedback control law that generates an optimal trajectory from any point in state-space without replanning.

Bayesian inference over admissible histories leads to irreversible kinetics.

problem Modeling irreversible processes in systems with uncertain histories.
method A Gibbs-type measure weighted by energy-dissipation action and observation constraints, interpreted as a Bayesian posterior.
result The measure concentrates on maximum-a-posteriori (MAP) histories, recovering classical deterministic evolution.

We study the effect of impairment on stochastic multi-armed bandits and develop new ways to mitigate it. Impairment effect is the phenomena where an agent only accrues reward for an action if they have played it at least a few times in the recent past. It is practically motivated by repetition and recency effects in do…

2018-11-22abs ↗pdf ↗

Suspensions of manifolds by circle surgeries are key in free action constructions.

problem Understanding free S1S^1-actions on smooth manifolds of dimension at least 3.
method Circle surgeries on S1imesMS^1 imes M yield suspensions Σ0MΣ_0M and Σ1MΣ_1M.
result Suspension operations ΣiΣ_i are fundamental in constructing and classifying manifolds with free S1S^1-actions.

Paper improves COCO problem, reducing constraint violation at the cost of slightly more regret.

problem Online Convex Optimization with adversarial constraints.
method Proposes new policies that trade off regret for reduced constraint violation.
result Achieves ildeO(dT+Tβ) ilde{O}(\sqrt{dT}+ T^β) regret and ildeO(dT1β) ilde{O}(dT^{1-β}) CCV.

New algorithm reduces constraint violation to O(T1/3)O(T^{1/3}) while maintaining O(T)O(\sqrt{T}) regret.

problem Minimizing static regret and cumulative constraint violation in constrained online convex optimization.
method Proposes an algorithm that achieves O(T)O(\sqrt{T}) regret and O(T1/3)O(T^{1/3}) cumulative constraint violation.
result Shows that O(T1/3)O(T^{1/3}) cumulative constraint violation is achievable with O(T)O(\sqrt{T}) regret.

Neural Index Policy for multi-action bandits with heterogeneous budgets.

problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.

New attacks on RL agents' action space improve understanding of cyber-physical systems vulnerabilities.

problem Understanding and improving the robustness of RL agents in CPS against action space attacks.
method Proposed white-box MAS and LAS attack algorithms to optimize and temporally couple attack budgets.
result LAS attacks cause significantly more performance degradation than MAS attacks.

For a Lie group G and a smooth manifold W, we study the difference between smooth actions of G on W and bundles over the classifying space of G with fiber W and structure group Diff(W). In particular, we exhibit smooth manifold bundles over BSU(2) that are not induced by an action. The main tool for reaching this goal …

2018-02-06abs ↗pdf ↗

Optimal algorithm for submodular maximization in distributed networked scenarios.

problem Maximizing submodular functions under distributed matroid constraints.
method Constraint-Distributed Continuous Greedy (CDCG) algorithm.
result Achieves a (11/e)(1-1/e) approximation factor, significantly outperforming sequential greedy methods.

We study continuous action reinforcement learning problems in which it is crucial that the agent interacts with the environment only through safe policies, i.e.,~policies that do not take the agent to undesirable situations. We formulate these problems as constrained Markov decision processes (CMDPs) and present safe p…

2019-01-28abs ↗pdf ↗

When the vacuum Einstein equations are cast in the form of hamiltonian evolution equations, the initial data lie in the cotangent bundle of the manifold MΣ of riemannian metrics on a Cauchy hypersurface Σ. As in every lagrangian field theory with symmetries, the initial data must satisfy constraints. But, unlike those …

2010-03-15abs ↗pdf ↗

New framework guides resource usage to achieve sublinear regret in adversarial settings.

problem Achieving sublinear regret in online decision making with changing reward and cost distributions.
method General primal-dual methods guided by spending plans that ensure balanced resource usage.
result Achieves sublinear regret with respect to spending plans that balance resource usage.

We analyse the definition of quasi-local energy in GR based on a Hamiltonian analysis of the Einstein-Hilbert action initiated by Brown-York. The role of the constraint equations, in particular the Hamiltonian constraint on the timelike boundary, neglected in previous studies, is emphasized here. We argue that a consis…

2010-08-25abs ↗pdf ↗

This paper examines a proposal for gauging non-linear sigma models with respect to a Lie algebroid action. The general conditions for gauging a non-linear sigma model with a set of involutive vector fields are given. We show that it is always possible to find a set of vector fields which will (locally) admit a Lie alge…

2019-05-02abs ↗pdf ↗