Paper presents a reduction-based framework for conservative bandits and RL with improved lower and upper bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new algorithm balances exploration and exploitation in online decision-making.
A new algorithm for contextual combinatorial bandits reduces regret.
New algorithm improves online learning under performance constraints.
Two new algorithms optimize rewards while respecting safety constraints in sequential decisions.
An online reinforcement learning algorithm is anytime if it does not need to know in advance the horizon T of the experiment. A well-known technique to obtain an anytime algorithm from any non-anytime algorithm is the "Doubling Trick". In the context of adversarial or stochastic multi-armed bandits, the performance of …
The paper tackles restless bandits with limited observation, proposing a method to analyze and approximate their optimal strategies.
Safety is a desirable property that can immensely increase the applicability of learning algorithms in real-world decision-making problems. It is much easier for a company to deploy an algorithm that is safe, i.e., guaranteed to perform at least as well as a baseline. In this paper, we study the issue of safety in cont…
Algorithm improves wildlife protection patrols.
Most bandit algorithm designs are purely theoretical. Therefore, they have strong regret guarantees, but also are often too conservative in practice. In this work, we pioneer the idea of algorithm design by minimizing the empirical Bayes regret, the average regret over problem instances sampled from a known distributio…
Develops algorithms for CCBs with non-linear costs, improving safety and performance.
Contextual bandit algorithms are essential for solving many real-world interactive machine learning problems. Despite multiple recent successes on statistically and computationally efficient methods, the practical behavior of these algorithms is still poorly understood. We leverage the availability of large numbers of …
New meta-learning approach for bandit policies that achieve high average reward.
Study optimal arms in combinatorial bandits with semi-bandit feedback and finite budget.
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
An algorithm for efficient experimentation in a dynamic environment with personalized preferences and context drifts.
Excessively changing policies in many real world scenarios is difficult, unethical, or expensive. After all, doctor guidelines, tax codes, and price lists can only be reprinted so often. We may thus want to only change a policy when it is probable that the change is beneficial. In cases that a policy is a threshold on …
A new framework estimates policy value robustly against confounders.
We study a novel multi-armed bandit problem that models the challenge faced by a company wishing to explore new strategies to maximize revenue whilst simultaneously maintaining their revenue above a fixed baseline, uniformly over time. While previous work addressed the problem under the weaker requirement of maintainin…
Sharp policy value estimation for contextual bandits with unobserved confounders.
Study best arm identification in restless bandits with unknown TPMs.
The paper addresses statistical inference issues in adaptive experiments.
New algorithm optimizes beam and rate allocation in mmWave systems for multiple users.
Theoretical analysis confirms non-conservative algorithms can converge to optimal policies.
Study finds conserved quantities for two types of curves on conformal sphere.
Survey on conservation laws for geometric PDEs.
The article discusses conservation laws for polyharmonic maps and their applications.
I consider the existence and structure of conservation laws for the general class of evolutionary scalar second-order differential equations with parabolic symbol. First I calculate the linearized characteristic cohomology for such equations. This provides an auxiliary differential equation satisfied by the conservatio…
Non-trivial conservation law found for a specific system.
This work connects symmetries and conserved quantities in machine learning.
Proposes a conservative exploration method for RL agents.
The conservation laws of the third order quasilinear scalar evolution equations are considered via differential system and characteristic cohomology. We find a subspace of 2 forms in the infinite prolonged space in which every conservation law has a unique representative. The structure of this subspace naturally gives …
Conservation law for weakly harmonic mappings in high dimensions.
Given a vector field on a manifold M, we define a globally conserved quantity to be a differential form whose Lie derivative is exact. Integrals of conserved quantities over suitable submanifolds are constant under time evolution, the Kelvin circulation theorem being a well-known special case. More generally, conserved…
In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent has to explore unknown actions, some of which can be bad, to learn better action…
We study higher-order conservation laws of the non-linearizable elliptic Poisson equation as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…
New conservation laws found for polyharmonic maps in critical dimension.
A challenging problem in complex networks is the network reconstruction problem from data. This work deals with a class of networks denoted as conserved networks, in which a flow associated with every edge and the flows are conserved at all non-source and non-sink nodes. We propose a novel polynomial time algorithm to …
MC-LSTM extends LSTM to conserve mass in neural networks.
The paper studies symmetries and conservation laws of non-diagonalisable hydrodynamic systems.
We present a connection between the Killing fields that arise in the loop-group approach to integrable systems and conservation laws viewed as elements of the characteristic cohomology. We use the connection to generate the complete set of conservation laws (as elements of the characteristic cohomology) for the Tzitzei…
We obtain necessary and sufficient conditions for the existence of "conservation laws" on null hypersurfaces for the wave equation on general four-dimensional Lorentzian manifolds. Examples of null hypersurfaces exhibiting such conservation laws include the standard null cones of Minkowski spacetime and the degenerate …
New neural network enforces mass conservation for better ice flow predictions.
The study finds resonance points in polarised curves with polynomial conserved quantities.
Following an approach of the second author for conformally invariant variational problems in two dimensions, we show in four dimensions the existence of a conservation law for fourth order systems, which includes both intrinsic and extrinsic biharmonic maps. With the help of this conservation law we prove the continuit…
There is a well-known example of integrable conservative system on , the case of Kovalevskaya in the dynamics of a rigid body, possessing an integral of fourth degree in momenta. Goryachev proposed a one-parameter family of examples of conservative systems on possessing an integral of fourth degree in moment…
Data symmetries in neural networks can generate conserved quantities.
RORL improves offline RL robustness with conservative smoothing.