Agents minimize individual regret in a networked multi-armed bandit problem.
problem Minimizing individual regret in a networked multi-armed bandit problem.
method Regret minimization algorithms for agents communicating over a network.
result Guaranteed individual expected regret of $\widetilde{O}\left(\sqrt{\left(1+\frac{K}{\left|\mathcal{N}\left(v
ight)
ight|}
ight)T}
ight)$ for each agent.
Improves bandit regret with small loss range, even with limited information.
problem Improving bandit regret with small loss range.
method Develops a novel technique to convert algorithms with regret depending on loss range to ones with regret depending only on effective range.
result Shows how to improve bandit regret guarantees with small loss range under certain assumptions.
New algorithm reduces regret in collaborative multi-agent bandit problems.
problem Optimizing decisions in a network of agents with communication delays.
method Follow-the-Regularized-Leader (FTRL) algorithm with suitable regularizers and communication protocols.
result Upper bound on individual regret matches lower bound up to a constant factor.
Multi-armed bandit problems are the most basic examples of sequential decision problems with an exploration-exploitation trade-off. This is the balance between staying with the option that gave highest payoffs in the past and exploring new options that might give higher payoffs in the future. Although the study of band…
We present and study a partial-information model of online learning, where a decision maker repeatedly chooses from a finite set of actions, and observes some subset of the associated losses. This naturally models several situations where the losses of different actions are related, and knowing the loss of one action p…
Meta-algorithm optimizes nonstochastic bandits with infinitely many experts.
problem Maximizing reward by choosing actions sequentially from a set of experts.
method Proposed a variant of Exp4.P for infinitely many experts and a meta-algorithm.
result Proved high-probability upper bound of i l d e O ( i ∗ K + K T ) ilde{\mathcal{O}} \big( i^*K + \sqrt{KT} \big) i l d e O ( i ∗ K + K T ) on regret. ABoB optimizes online configuration tuning by clustering parameters and accelerating learning.
problem Online optimization in large, dynamic parameter spaces.
method Hierarchical adversarial bandit framework.
result Significant performance gains in adversarial metric scenarios.
New dynamic allocation methods for multi-armed bandit models.
problem Dynamic allocation problems in multi-armed bandit models.
method New types of dynamic allocation problems and proofs for Gittins index decomposition.
result New proofs for Gittins index decomposition and related results.
PAC-Bayesian analysis improves lifelong learning in multi-armed bandits.
problem Improving lifelong learning in multi-armed bandits.
method PAC-Bayesian analysis for deriving lower bounds and proposing lifelong learning algorithms.
result Proposed algorithms outperform baseline methods in lifelong multi-armed bandit problems.
BaSE policy optimizes multi-armed bandits with batched data.
problem Optimizing multi-armed bandits with batched data.
method BaSE (batched successive elimination) policy for batched multi-armed bandits.
result Achieves rate-optimal regrets with adaptive batch sizes.
Paper tackles risk-aware portfolio selection using multi-armed bandit.
problem Sequential portfolio selection under uncertainty.
method Incorporates risk-awareness into multi-armed bandit, constructs portfolio through asset filtering and risk minimization.
result Achieves balance between risk and return.
OSOM solves multi-armed and linear contextual bandits efficiently.
problem Simultaneously optimal algorithm for multi-armed and linear contextual bandits.
method Design of a single computationally efficient algorithm that adapts to both regimes.
result Simultaneously optimal regret rates in both simple multi-armed and linear contextual bandits.
Survey of multi-armed bandit applications in various fields.
problem Optimizing decisions with limited feedback in diverse contexts.
method Comprehensive review of recent developments and trends.
result Identification of important current trends and future directions.
Study uses randomized allocation for delayed rewards in multi-armed bandits.
problem Delayed rewards in contextual multi-armed bandits.
method Randomized allocation with nonparametric estimation.
result Strongly consistent strategy for delayed rewards.
Adapts multi-armed bandits to contextual bandits using logistic regression.
problem Online decision-making with contextual information and binary rewards.
method Adapts multi-armed bandits policies to contextual bandits using logistic regression and bootstrapping.
result Adaptive-Greedy algorithm shows better performance than upper confidence bound and Thompson sampling.
Paper uses subjective logic to estimate uncertainty in multi-armed bandit problems.
problem Estimating uncertainty in multi-armed bandit problems.
method Formalism of subjective logic applied to multi-armed bandits, proposing new algorithms.
result Subjective logic quantities enable useful assessment of uncertainty.
Paper simplifies Gittins indices calculation for bandits.
problem Difficulty in calculating Gittins indices for multi-armed bandits.
method Accessible general methodology for calculating Gittins indices.
result Removes computation barrier for Gittins indices.
New algorithm balances exploration and exploitation in multi-armed bandits with structured priors.
problem Balancing exploration and exploitation in multi-armed bandits with structured priors.
method Value-function-driven online planning techniques with n-step lookahead.
result Sub-linear performance guarantee and strong practical performance in structured priors.
Algorithm improves multi-armed bandit performance by transferring reward samples.
problem Sequential multi-armed bandit problem with changing reward distributions.
method UCB algorithm with reward sample transfer.
result Significant improvement in cumulative regret over standard UCB.
Quantum algorithms for multi-armed bandits are explored with limited reward access.
problem Exploring quantum speed-ups in multi-armed bandit problems with limited reward information.
method Introduced new bandit models and showed query complexity equivalence with classical algorithms.
result No quadratic speed-up is possible for multi-armed bandits with limited reward access.
Study on Pareto optimality in multi-objective bandit problems.
problem Pareto optimality in multi-objective multi-armed bandit problems.
method Formulated adversarial multi-objective multi-armed bandit, defined Pareto regrets, presented algorithms, established upper and lower bounds.
result New algorithms are optimal in adversarial settings and nearly optimal in stochastic settings.
Unified formulation bridges adversarial and nonstationary bandits.
problem Handling time-varying reward distributions in multi-armed bandit problems.
method Unified oracle that switches between adversarial and nonstationary bandit oracles based on window size.
result Optimal regret achieved with matching lower bound.
In this paper we propose a multi-armed bandit inspired, pool based active learning algorithm for the problem of binary classification. By carefully constructing an analogy between active learning and multi-armed bandits, we utilize ideas such as lower confidence bounds, and self-concordant regularization from the multi…
The paper tackles lifelong learning in multi-armed bandits, aiming to minimize average regret over multiple tasks.
problem Minimizing average regret in multi-armed bandits over multiple tasks.
method Confidence interval tuning of UCB algorithms and greedy algorithms applied to a bandit over bandit approach.
result Empirical improvement over previous work in the mortal bandit problem.
Study finds optimal regret bound for multi-armed bandit problem with expert advice.
problem Optimizing decision-making in a multi-armed bandit problem with expert advice.
method Proved a tight lower bound matching the upper bound of Kale (2014) for minimax expected regret.
result The minimax optimal expected regret is Θ(√(T K log (N/K))) for the problem.
New method for identifying best arm in batched multi-armed bandit problems.
problem Identifying the best arm in multi-armed bandit problems where arms are sampled in batches.
method General linear programming framework for best arm identification in batched multi-armed bandit problems.
result Demonstrated good performance in numerical studies compared to UCB-type or Thompson sampling methods.
Study collaborative learning among multi-agents in multi-armed bandits.
problem Minimizing group cumulative regret in a heterogeneous multi-agent setting.
method Developed decentralized algorithms for collaboration between N N N agents learning M M M stochastic multi-armed bandits. result Proved near-optimal behavior of proposed algorithms for group regret.
The paper explores sampling problems and shows minimal exploration is needed.
problem Exploration-exploitation trade-off in sampling.
method Systematic definition of regret, proposal of a simple algorithm.
result Near-optimal regret bounds achieved with minimal exploration.
New algorithm for nonstationary multi-armed bandits with optimal performance.
problem Nonstationary multi-armed bandits with changing model parameters over time.
method Adaptive Resetting Bandit (ADR-bandit) algorithm using adaptive windowing techniques.
result ADR-bandit achieves nearly optimal performance in both abrupt and gradual changes.
Unified DP framework for multi-armed bandits with privacy cost.
problem Privacy in multi-armed bandit algorithms.
method Differential privacy framework, unified graphical model, lower bounds derivation.
result Regret is increased by a factor dependent on privacy level ε.
Unified approach to correlated multi-armed bandits reduces regret significantly.
problem Correlated rewards in multi-armed bandits.
method Developed a unified approach to leverage reward correlations and presented algorithms with rigorous analysis.
result C-UCB algorithm pulls non-competitive arms only O(1) times, improving over classic algorithms.
Paper tackles transfer learning for contextual multi-armed bandits under covariate shift.
problem Nonparametric contextual multi-armed bandits with covariate shift.
method Established minimax rate of convergence, proposed transfer learning algorithm.
result Achieved near-optimal statistical guarantees for learning in target domain.
Proposes Genetic Thompson Sampling for multi-armed bandits, improving performance in nonstationary settings.
problem Improving sequential decision making tasks of online learning agents using multi-armed bandits.
method Integrates genetic principles into Thompson Sampling for multi-armed bandits.
result Significantly outperforms baselines in nonstationary settings.
We present a formal model of human decision-making in explore-exploit tasks using the context of multi-armed bandit problems, where the decision-maker must choose among multiple options with uncertain rewards. We address the standard multi-armed bandit problem, the multi-armed bandit problem with transition costs, and …
New method for contextual bandits with corrupted context.
problem Contextual bandits with corrupted context in online settings.
method Combining contextual bandit and multi-armed bandit approaches.
result Improved learning from all iterations, including corrupted ones.
This work optimizes marketing by targeting persuadable customers with causal effects.
problem Optimizing marketing ROI by targeting only those who would be influenced.
method Causal contextual multi-armed bandits, incorporating causal inference and uplift modeling.
result Preliminary experiments show the approach improves marketing ROI.
A new algorithm solves a regional multi-armed bandit problem with group information.
problem Optimizing decisions with unknown parameters across groups.
method UCB-g algorithm combining UCB and greedy principles.
result Proves the order-optimality of UCB-g and establishes a matching lower bound.
Regularized contextual bandits use bins to solve multi-armed bandit problems.
problem Contextual bandit problems with a known baseline policy.
method Nonparametric model, splitting context space into bins, solving bandit instances independently.
result Intermediate convergence rates interpolating between slow and fast rates.
New algorithm prevents strategic replication in multi-armed bandit problems.
problem Strategic replication by agents can exploit bandit algorithms' balance.
method Designs Hierarchical UCB (H-UCB) and Robust Hierarchical UCB (RH-UCB) algorithms.
result Achieves O ( ln T ) O(\ln T) O ( ln T ) -regret and sublinear regret in realistic scenarios. A new framework detects changes in multi-armed bandit problems.
problem Change in reward distributions over time in multi-armed bandit problems.
method Change-detection (CD) based UCB policies, CUSUM-UCB, PHT-UCB.
result CUSUM-UCB obtains the best known regret upper bound.
A version of indifference valuation of a European call option is proposed that includes statistical regularities of nonstochastic randomness. Classical relations (forward contract value and Black-Scholes formula) are obtained as particular cases. We show that in the general case of nonstochastic randomness the minimal …
This paper applies Thompson Sampling to asymmetric α \alpha α -stable bandits for financial and wireless data.
problem Optimizing exploration-exploitation in multi-armed bandits with asymmetric α \alpha α -stable distributions. method Thompson Sampling applied to unknown asymmetric α \alpha α -stable reward distributions. result Demonstrates effectiveness of Thompson Sampling for asymmetric α \alpha α -stable bandits. Deep learning tackles contextual multi-armed bandits with principled exploration.
problem Contextual multi-armed bandits in industrial applications.
method Bayesian neural network with dropout for non-linear modeling and Thompson sampling for principled exploration.
result Substantially reduces regret compared to existing methods.
The paper examines how loss aversion impacts multi-armed bandit decisions over long periods.
problem The impact of loss aversion on multi-armed bandit decisions over long periods.
method A new central limit theorem for measures with history-dependent variances, derived under risk aversion in gains and risk loving in losses.
result Consequences of loss aversion for asymptotic properties are derived in analytical results.
Two non-communicating players minimize regret in a multi-armed bandit game.
problem Optimal regret in non-communicating multi-armed bandit players.
method Proposed a strategy with no collisions, achieving near-optimal regret.
result Near-optimal regret of O ( T log ( T ) ) O(\sqrt{T \log(T)}) O ( T log ( T ) ) with very high probability. The study uses a multi-armed bandit model to analyze and mitigate hiring discrimination.
problem Hiring discrimination due to insufficient data on worker skill and characteristics.
method Multi-armed bandit model to simulate firms' learning process and policy solutions.
result Temporary affirmative actions effectively alleviate discrimination caused by data insufficiency.
Study optimal allocation in uncertain multi-armed bandits using Gittins' theorem.
problem Optimal allocation in uncertain multi-armed bandits.
method Theoretical analysis based on nonlinear expectations, with relaxation in optimality definition.
result Gittins' allocation index provides optimal choices under strong independence and relaxed optimality conditions.
Optimal strategy found for constrained multi-armed bandit problems.
problem Constrained multi-armed bandit problems.
method Extended ε_t-greedy strategy with asymptotic optimality.
result Asymptotic optimality achieved with a simple strategy.