New method identifies Condorcet winner in dueling bandits with improved sample complexity.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper tackles combinatorial pure exploration for dueling bandits, aiming to find the best candidate-position match.
New algorithm reduces dynamic regret in non-stationary dueling bandits using a weighted Borda score.
New algorithms improve dueling bandit performance in multiplayer settings.
Improved algorithm for adaptive dueling bandits with near-optimal regret bound.
New algorithms for batched dueling bandits with improved regret bounds.
Algorithm minimizes regret in non-stationary dueling bandits with unknown parameters.
Unified framework for best arm identification and dueling bandits regret minimization.
New algorithm optimizes dueling bandits for both stochastic and adversarial preferences.
New algorithm for identifying Condorcet team in noisy comparisons.
Unified framework for ranking-and-selection with multiple correct answers and non-answerable estimates