The paper analyzes the sliding regret of stochastic bandit algorithms.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms achieve optimal regret in sliding window model with limited memory.
We study the multi-player stochastic multiarmed bandit (MAB) problem in an abruptly changing environment. We consider a collision model in which a player receives reward at an arm if it is the only player to select the arm. We design two novel algorithms, namely, Round-Robin Sliding-Window Upper Confidence Bound\# (RR-…
New TS algorithms improve performance in non-stationary multi-armed bandit problems.
Study tackles infinitely many-armed bandits with rotting rewards, achieving tight regret bounds.
We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem setting, we propose an algorithm using a sliding window approach and provide performance guarantees for the regret evaluated against the op…
New algorithm tackles non-stationary reinforcement learning with general function approximation.
We study the non-stationary stochastic multiarmed bandit (MAB) problem and propose two generic algorithms, namely, the limited memory deterministic sequencing of exploration and exploitation (LM-DSEE) and the Sliding-Window Upper Confidence Bound# (SW-UCB#). We rigorously analyze these algorithms in abruptly-changing a…
We introduce algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary linear stochastic bandit setting. It captures natural applications such as dynamic pricing and ads allocation in a changing environment. We show how the difficulty posed by the non-stationarity can be overcome by a nov…
We propose a sliding surface for systems on the Lie group . The sliding surface is shown to be a Lie subgroup. The reduced-order dynamics along the sliding subgroup have an almost globally asymptotically stable equilibrium. The sliding surface is used to design a sliding-mode controller for t…
New algorithm optimizes resource allocation in non-stationary networks.
New TS-like algorithms tackle rising bandits with improved regret analysis.
KeRNS tackles non-stationary reinforcement learning in metric spaces.
New RL method tackles dynamic MDPs with evolving rewards and states.
Presentations for involutions on non-orientable surfaces up to genus 5.
Study local exploration on dynamic graphs with time-varying edges.
We present a new operation to be performed on elements in a Garside group, called cyclic sliding, which is introduced to replace the well known cycling and decycling operations. Cyclic sliding appears to be a more natural choice, simplifying the algorithms concerning conjugacy in Garside groups and having nicer theoret…
Novel time series forecasting method using sliding window signatures.
Proposes a sliding window method for better portfolio trading.
New algorithms for GLMs adapt to non-stationary contexts.
This paper refines the weighted strategy for non-stationary parametric bandits, improving regret bounds.
Paper simplifies proof of slide-equivalence in crown diagrams.
Algorithm minimizes control regret for non-stationary LQR systems.
Optimizes sliding window approach for tracking Gaussian densities.
Let denote a closed nonorientable surface of genus . For the mapping class group is generated by Dehn twists and one crosscap slide (-homeomorphism) or by Dehn twists and a crosscap transposition. Margalit and Schleimer observed that Dehn twists have nontrivial roots. We gi…
Small bubbles sliding on a boundary maintain half-spherical shape.
We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applications such as advertisement allocation, dynamic pricing, and traffic network routing in changing environments. We show how the difficulty posed…
We show that reducible braids which are, in a Garside-theoretical sense, as simple as possible within their conjugacy class, are also as simple as possible in a geometric sense. More precisely, if a braid belongs to a certain subset of its conjugacy class which we call the stabilized set of sliding circuits, and if it …
A new method optimizes in nonstationary environments with many arms efficiently.
This paper refines the weighted strategy for non-stationary parametric bandits and MDPs, improving regret bounds.
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed to evolve over time, as long as their respective total variations, quantified by suitable metrics, do not exceed certain variation budgets. …
If a variational problem comes with no boundary conditions prescribed beforehand, and yet these arise as a consequence of the variation process itself, we speak of a free boundary values variational problem. Such is, for instance, the problem of finding the shortest curve whose endpoints can slide along two prescribed …
AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.
A theorem of Kirby states that two framed links in the 3-sphere produce orientation-preserving homeomorphic results of surgery if they are related by a sequence of stabilization and handle-slide moves. The purpose of the present paper is twofold: First, we give a sufficient condition for a sequence of handle-slides on …
Kirby color defined in Khovanov homology for 4D handlebodies.
A new method SLIDE ensures fairness in AI models.
We analyze higher-dimensional sliding puzzles, finding solvability patterns.
Crosscap slide is a homeomorphism of a nonorientable surface of genus at least 2, which was introduced under the name Y-homeomorphism by Lickorish as an example of an element of the mapping class group which cannot be expressed as a product of Dehn twists. We prove that the subgroup of the mapping class group of a clos…
We consider the problem of designing an allocation rule or an "online learning algorithm" for a class of bandit problems in which the set of control actions available at each time is a convex, compact subset of . Upon choosing an action at time , the algorithm obtains a noisy value of the unkno…
We present a novel approach to train pixel resolution segmentation models on whole slide images in a weakly supervised setup. The model is trained to classify patches extracted from slides. This leads the training to be made under noisy labeled data. We solve the problem with two complementary strategies. First, the pa…
In many applications, monitoring area under the ROC curve (AUC) in a sliding window over a data stream is a natural way of detecting changes in the system. The drawback is that computing AUC in a sliding window is expensive, especially if the window size is large and the data flow is significant. In this paper we propo…
We propose a means by which some categorifications can be evaluated at a root of unity. This is implemented using a suitable localization in the context of prior work by the authors on categorification of the Jones-Wenzl projectors. Within this construction we define objects, invariant under handle slides, which decate…
TAKDE optimizes kernel density estimation for real-time dynamic processes.
Bedside monitors in Intensive Care Units (ICUs) frequently sound incorrectly, slowing response times and desensitising nurses to alarms (Chambrin, 2001), causing true alarms to be missed (Hug et al., 2011). We compare sliding window predictors with recurrent predictors to classify patient state-of-health from ICU multi…
Study predicts cryptocurrency trends using LSTM model.
A new method for real-time CCA on streaming data.
New algorithm for nonstationary GLBs reduces computation and memory costs.
New robustness certificates for streaming models with a sliding window.