Paper develops bandit algorithms for nonstationary nonconvex optimization.
problem Nonstationary online nonconvex optimization problems.
method Proposes and analyzes bandit algorithms for nonconvex functions with nonstationary regret.
result Develops bandit versions of Newton's method for nonstationary nonconvex optimization.
Study risk-sensitive reinforcement learning with Lipschitz dynamic risk measures, establishing regret bounds.
problem Risk-sensitive reinforcement learning in Markov decision processes.
method Two model-based algorithms for Lipschitz dynamic risk measures, focusing on regret bounds.
result Upper bounds demonstrate optimal dependencies on actions and episodes, reflecting risk sensitivity vs. sample complexity trade-off.
Protocol learns pure quantum states with minimal disturbance.
problem Efficiently learn quantum states with minimal disturbance.
method Sequential measurements with minimal disturbance.
result Achieves maximal precision with polylogarithmic regret.
Paper proposes algorithms to minimize both dynamic and adaptive regret simultaneously.
problem Traditional regret minimization algorithms are suboptimal for changing environments.
method Developed novel online algorithms to minimize dynamic and adaptive regret simultaneously.
result Proposed algorithms minimize dynamic and adaptive regret over any interval.
New algorithms reduce regret in online MDPs by adapting to data and variance.
problem Adapting to both adversarial and stochastic environments in online MDPs.
method Develops algorithms based on global optimization and policy optimization, using optimistic follow-the-regularized-leader with log-barrier regularization.
result Achieves refined data-dependent and variance-dependent regret bounds.
The paper analyzes the sliding regret of stochastic bandit algorithms.
problem Measuring the one-shot behavior of no-regret algorithms in stochastic bandits.
method Introducing sliding regret to measure the worst pseudo-regret over a time-window.
result Randomized methods have optimal sliding regret, while index policies have the worst possible sliding regret.
Paper introduces a new G ⋆ G^\star G ⋆ regret measure for online convex optimization with smooth losses.
problem Online convex optimization with smooth losses.
method Introduces a new G ⋆ G^\star G ⋆ regret measure that depends on the cumulative squared gradient norm. result The G ⋆ G^\star G ⋆ regret can be arbitrarily sharper than existing measures when losses have vanishing curvature. New algorithms reduce dynamic regret in online MDPs with changing losses.
problem Online MDPs with adversarial loss changes and known transitions.
method Dynamic regret measure, novel ensemble algorithms for three models.
result Provably optimal dynamic regret bounds for episodic SSP, improved bounds for predictable environments.
New measure of policy regret shows compatibility with traditional external regret in adversarial games.
problem Incompatibility between traditional and new policy regret measures in adaptive adversaries.
method Revisited policy regret and compared it with external regret; introduced policy equilibrium.
result Policy regret and external regret are compatible in adversarial games.
The paper derives a new theorem for predicting batches of data.
problem Finding lower bounds on minimal batch regret.
method Derives a conditional version of the regret-capacity theorem.
result Reveals a connection between conditional Rényi divergence and conditional Sibson's mutual information.
New algorithm reduces dynamic regret for MDPs with unknown transition and adversarial rewards.
problem Episodic linear mixture MDPs with unknown transition and adversarial rewards.
method Combines occupancy-measure-based global optimization and policy-based variance-aware value-targeted regression.
result Achieves near-optimal dynamic regret of O ~ ( d H 3 K + H K ( H + P ˉ K ) ) \widetilde{\mathcal{O}}(d \sqrt{H^3 K} + \sqrt{HK(H + \bar{P}_K)}) O ( d H 3 K + H K ( H + P ˉ K ) ) . Novel Orlicz regrets consistently bound environmental variable statistics.
problem Consistent evaluation of stochastic environmental variables like water quality indices.
method Proposed novel Orlicz regrets for upper and lower bounds.
result Explicit linkage between Orlicz regrets and divergence risk measures.
New algorithm minimizes Bayesian regret in offline linear bandits.
problem Minimizing Bayesian regret in offline linear bandits.
method Proposes a new algorithm that directly minimizes upper bounds on Bayesian regret using conic optimization.
result Upper bounds are tight and guarantee superior performance compared to LCB.
New algorithm achieves both static and dynamic regret optimally against an oblivious adversary for deterministic losses.
problem Achieving optimal static and dynamic regret simultaneously in adversarial bandits.
method Extends impossibility result to deterministic losses, uses negative static regret and Blackwell approachability.
result First algorithm achieving optimal static and dynamic regret simultaneously against an oblivious adversary.
New calibration measure SCDL improves trust in AI predictions.
problem Improving trust in AI predictions by ensuring they are both actionable and testable.
method Introducing SCDL, a new calibration measure that is fully actionable and testable.
result SCDL is the first calibration measure that is fully actionable and testable.
New definition of regret for nonconvex online learning models.
problem Intractability of standard regret measures for nonconvex models.
method Introduced a local gradient based regret definition.
result Our definition provides more interpretable bounds for forecasting.
Algorithm learns from changing zero-sum games with no regret.
problem Learning in time-varying zero-sum games.
method Developed a single parameter-free algorithm with three performance measures.
result Algorithm recovers best known results for fixed games and adapts to non-stationarity.
We consider a setting where an agent's uncertainty is represented by a set of probability measures, rather than a single measure. Measure-bymeasure updating of such a set of measures upon acquiring new information is well-known to suffer from problems; agents are not always able to learn appropriately. To deal with the…
New algorithm reduces prediction error in online learning without knowing base measure.
problem Smoothed online learning without knowledge of base measure.
method R-Cover algorithm based on recursive coverings.
result First algorithm to guarantee sublinear regret for agnostic smoothed online learning without prior knowledge of base measure.
Paper improves adaptive regret for convex and smooth functions.
problem Online convex optimization in changing environments.
method Develops adaptive algorithms exploiting both convexity and smoothness.
result Regret bounds are comparable to worst-case results but tighter when comparators have small losses.
New algorithms reduce contextual bandits' regret without knowing reward noise variances.
problem Reducing regret in contextual bandits with unknown reward noise variances.
method Developed new algorithms based on the optimism principle.
result Regret scales as the square root of the sum of measurement variances, not the time horizon.
Paper proposes new γ γ γ -regret measure for non-episodic RL.
problem Measuring performance in non-episodic RL environments.
method Introduces γ γ γ -regret as a new performance measure and derives bounds. result Closed the gap between lower and upper bounds for γ γ γ -regret. ERTS uses Thompson sampling for Gaussian entropic risk bandits, achieving regret bounds.
problem Risk in decision making complicates reward maximization in MAB problems.
method ERTS (Entropic Risk Thompson Sampling) using Thompson sampling with an entropic risk measure.
result Regret bounds for ERTS under entropic risk measure provided.
Paper tackles online learning with interval regret, achieving adaptive bounds.
problem Non-stationary online learning over time intervals.
method Two-layer online ensemble structure with gradient variation.
result Achieves strong theoretical guarantees with adaptive bounds.
Online learning algorithms are designed to learn even when their input is generated by an adversary. The widely-accepted formal definition of an online algorithm's ability to learn is the game-theoretic notion of regret. We argue that the standard definition of regret becomes inadequate if the adversary is allowed to a…
OE2D framework reduces contextual bandits to offline regression for near-optimal regret.
problem Efficiently learning contextual bandits with large action spaces and complex reward functions.
method Offline Estimation to Decisions (OE2D) algorithm that reduces contextual bandits to offline regression.
result Near-optimal regret for contextual bandits with large action spaces and O ( log T ) O(\log T) O ( log T ) calls to an offline regression oracle. The paper sets lower bounds on regret for optimizing noisy Gaussian processes.
problem Sequentially optimizing a black-box function with noisy samples and bandit feedback.
method Algorithm-independent lower bounds on simple and cumulative regret for Gaussian process bandit optimization.
result Lower bounds on simple and cumulative regret for various kernels, matching existing upper bounds.
A new policy for contextual bandits adapts to reward vector shifts.
problem Learning under reward vector shifts with ordered rewards.
method Adaptive-discretization and optimistic elimination policy.
result Established upper bounds on preference-based regret.
EBUCB framework achieves optimal regret with bounded approximate inference error.
problem Theoretical gap between practical performance and theoretical justification of Bayesian bandit algorithms with approximate inference.
method Enhanced Bayesian Upper Confidence Bound (EBUCB) framework that accommodates bandit problems with approximate inference.
result EBUCB achieves optimal regret order O ( log T ) O(\log T) O ( log T ) under certain conditions on inference error. This paper studies risk-averse online learning, showing differences from risk-neutral approaches.
problem Risk-averse online learning under mean-variance performance measure.
method Analyzes bandit and full information settings, establishes fundamental limitations.
result Worst-case regret is lower bounded by Ω ( T ) Ω(T) Ω ( T ) , contrasting with Ω ( T ) Ω(\sqrt{T}) Ω ( T ) for risk-neutral learning. UCRL-CMDP algorithm optimizes RL with constraints on average costs.
problem Optimizing RL in MDPs with average cost constraints.
method Model-based RL algorithms maximizing reward while keeping costs within bounds.
result UCRL-CMDP algorithm's expected regret is upper-bounded by $T^{2\slash 3}$ .
The paper develops algorithms to minimize risk and regret in uncertain decisions.
problem Minimizing risk and regret in multistage decisions under uncertainty.
method Established dual representations and used Lagrangian duality theory to develop progressive hedging algorithms.
result Modified progressive hedging algorithm can handle new linkage constraints.
Framework reduces contextual bandit learning to offline regression with near-optimal regret.
problem Efficient learning with large action spaces and complex reward functions.
method Offline Estimation to Decisions (OE2D) algorithm that minimizes regret with near-optimal oracle calls.
result Near-optimal regret for contextual bandits with large action spaces and O ( l o g ( T ) ) O(log(T)) O ( l o g ( T )) offline oracle calls. Consider the classical problem of predicting the next bit in a sequence of bits. A standard performance measure is {\em regret} (loss in payoff) with respect to a set of experts. For example if we measure performance with respect to two constant experts one that always predicts 0's and another that always predicts 1's …
Algorithm minimizes regret while adhering to unknown safety constraints.
problem Online learning with unknown safety constraints.
method General meta-algorithm leveraging online regression and learning oracles.
result Concrete algorithm with T \sqrt{T} T regret for linear constraints. Study bandit problems with BMO functions, achieving poly-log δ δ δ -regret.
problem Bandit problems with discontinuous, unbounded functions.
method Developed a toolset for BMO bandits and an algorithm achieving poly-log δ δ δ -regret. result Achieved poly-log δ δ δ -regret against an optimal arm after removing a δ δ δ -sized portion of the arm space. New algorithm reduces dynamic regret by adapting to comparator complexity.
problem Nonstationary sequential decision making with unbounded domains.
method Sparse coding framework to adapt to comparator complexity.
result Improves dynamic regret bounds by adapting to comparator energy and sparsity.
Paper characterizes minimax regret rates for online ranking with top-k feedback.
problem Analyzing online ranking with partial feedback.
method Developed techniques from partial monitoring to characterize minimax regret rates.
result Full characterization of minimax regret rates for Precision@n.
New betting strategy reduces regret to ln(ln n) with protection against adversarial data.
problem Tackles the problem of minimizing regret in betting against adversarial and stochastic data.
method Combines insights from Robbins and Cover, using a mixture strategy.
result Exhibits a regret of O(ln(ln n)) on almost all paths, with O(log n) regret on the complement.
The paper introduces new measures to quantify variability in decision tree models due to observational multiplicity.
problem The variability in decision tree models due to observational multiplicity.
method Introduces leaf regret and structural regret to decompose observational multiplicity.
result Structural regret is the primary driver of observational multiplicity, accounting for over 15 times the variability of leaf regret in some datasets.
New robust bandit algorithm for clinical trials reduces sensitivity to outlier data.
problem Adaptive clinical trials need a robust bandit algorithm to handle outlier data.
method Proposes a new robustness criterion and modifies BESA algorithm for bandit problems.
result Empirical evaluation shows improved performance compared to standard bandit algorithms.
Near-optimal regret in distributed bandit learning with efficient communication protocols.
problem Minimizing total regret in collaborative bandit learning with limited communication.
method Proposed communication protocols for distributed multi-armed and linear bandits with near-optimal regret and efficient communication costs.
result Achieved near-optimal regret with communication costs independent of time horizon and number of arms.
New algorithms and bounds for contextual bandits using surrogate losses.
problem Efficiently solving contextual bandit problems with margin-based regret bounds.
method Use of surrogate losses (ramp and hinge) to derive new regret bounds and algorithms.
result Derives new margin-based regret bounds and efficient algorithms for contextual bandits.
Efficient methods reduce projections in non-stationary online learning.
problem Optimizing dynamic and adaptive regret in non-stationary online learning environments.
method Presented efficient methods reducing the number of projections per round from O ( log T ) O(\log T) O ( log T ) to 1 1 1 . result Reduced number of projections per round from O ( log T ) O(\log T) O ( log T ) to 1 1 1 for optimizing dynamic and adaptive regret. This paper considers the stability of online learning algorithms and its implications for learnability (bounded regret). We introduce a novel quantity called {\em forward regret} that intuitively measures how good an online learning algorithm is if it is allowed a one-step look-ahead into the future. We show that given…
New online conformal prediction methods minimize strongly adaptive regret and achieve near-optimal coverage.
problem Uncertainty quantification in online settings with changing data distributions.
method Developed new online conformal prediction methods that minimize strongly adaptive regret.
result Achieve near-optimal strongly adaptive regret and approximately valid coverage.
The paper explores dynamic regret with switching cost in online decision making.
problem The relation between dynamic regret and switching cost in online decision making.
method Investigates two classic online settings: Online Algorithms (OA) and Online Convex Optimization (OCO). Provides a new theoretical analysis framework.
result The switching cost impacts dynamic regret differently in OA and has no impact in OCO.
New algorithms reduce risk in reinforcement learning with provable regret bounds.
problem Risk-sensitive reinforcement learning in Markov decision processes.
method Two novel DRL algorithms leveraging the independence property of entropic risk measure.
result Regret bounds of i l d e O ( exp ( ∣ β ∣ H ) − 1 ∣ β ∣ H S 2 A K ) ilde{\mathcal{O}}(\frac{\exp(|β| H)-1}{|β|}H\sqrt{S^2AK}) i l d e O ( ∣ β ∣ e x p ( ∣ β ∣ H ) − 1 H S 2 A K ) for model-free and model-based algorithms.