KL-UCB-switch optimizes bandit strategies for both distribution-dependent and distribution-free performance.
problem Optimizing regret bounds for stochastic bandits.
method Combining MOSS and KL-UCB strategies.
result Achieves both optimal distribution-dependent and distribution-free regret bounds.
KL-UCB+ policy outperforms KL-UCB empirically in stochastic bandits.
problem Optimizing decisions in a stochastic bandit problem.
method Demonstrates a simple proof of asymptotic optimality for KL-UCB+ policy.
result KL-UCB+ policy achieves asymptotically optimal regret bound.
New algorithm outperforms existing ones in multi-player bandit problems without sensing.
problem Decentralized multi-player multi-armed bandit problem without collision or sensing info.
method Randomized Selfish KL-UCB, inspired by Selfish KL-UCB, with low complexity.
result Randomized Selfish KL-UCB outperforms state-of-the-art algorithms in almost all environments.
Adaptive KL-UCB algorithm for Markov and i.i.d. rewards.
problem Regret minimization for Markovian and i.i.d. rewards in MAB problems.
method Identifies Markovian vs. i.i.d. rewards, switches between KL-UCB variants.
result Logarithmic regret for both i.i.d. and Markovian settings.
UCBoost improves bandit algorithms to balance optimality and complexity.
problem Finding near-optimal multi-armed bandit algorithms with low complexity.
method Boosting approach to Upper Confidence Bound (UCB) algorithms.
result UCBoost algorithms achieve near-optimal regret guarantees with significantly reduced computational complexity.
New findings show many popular bandit algorithms are unstable, contradicting minimax optimality.
problem Challenges in statistical inference from bandit algorithms due to adaptive, non-i.i.d. nature.
method Analysis of stability properties of optimism-based bandit algorithms.
result Widely used minimax-optimal UCB-style algorithms are unstable.
This paper is about index policies for minimizing (frequentist) regret in a stochastic multi-armed bandit model, inspired by a Bayesian view on the problem. Our main contribution is to prove that the Bayes-UCB algorithm, which relies on quantiles of posterior distributions, is asymptotically optimal when the reward dis…
Proposes k-NN UCB for multi-armed bandits with covariates.
problem Optimizing decisions in multi-armed bandits with covariates.
method k-Nearest Neighbour UCB algorithm for low-dimensional data.
result Minimax optimal regret bound and empirical advantage.
The profitable bandit problem aims to maximize earnings by choosing actions with uncertain rewards.
problem Maximizing earnings from uncertain rewards in a subset selection problem.
method Adapted and studied three strategies: kl-UCB, Bayes-UCB, and Thompson Sampling.
result Establishes asymptotic optimality for each strategy with finite time regret bounds.
The exploration/exploitation (E/E) dilemma arises naturally in many subfields of Science. Multi-armed bandit problems formalize this dilemma in its canonical form. Most current research in this field focuses on generic solutions that can be applied to a wide range of problems. However, in practice, it is often the case…
GLR-klUCB detects change-points in non-stationary bandits efficiently.
problem Non-stationary bandit problems with piecewise stationary behavior.
method Combines kl-UCB with a changepoint detector based on GLR.
result Achieves O ( T A Υ T log ( T ) ) O(\sqrt{TA Υ_T\log(T)}) O ( T A Υ T log ( T ) ) regret for some instances. The paper tackles best arm identification with minimal regret in experiments.
problem Identifying the best arm with minimal regret in experiments.
method Information-theoretic techniques and Double KL-UCB algorithm.
result Achieves asymptotic optimality in identifying the best arm with minimal regret.
We propose the kl-UCB ++ algorithm for regret minimization in stochastic bandit models with exponential families of distributions. We prove that it is simultaneously asymptotically optimal (in the sense of Lai and Robbins' lower bound) and minimax optimal. This is the first algorithm proved to enjoy these two propertie…
We study a generalization of the multi-armed bandit problem with multiple plays where there is a cost associated with pulling each arm and the agent has a budget at each time that dictates how much she can expect to spend. We derive an asymptotic regret lower bound for any uniformly efficient algorithm in our setting. …
The paper analyzes the sliding regret of stochastic bandit algorithms.
problem Measuring the one-shot behavior of no-regret algorithms in stochastic bandits.
method Introducing sliding regret to measure the worst pseudo-regret over a time-window.
result Randomized methods have optimal sliding regret, while index policies have the worst possible sliding regret.
A new algorithm for better decision-making in recommendation systems.
problem Stochastic multi-armed bandit problem and cold start problem in recommender systems.
method Proposes Hellinger-UCB, a variant of UCB algorithm using squared Hellinger distance.
result Hellinger-UCB reaches the theoretical lower bound and outperforms other algorithms in practical applications.
New algorithms improve privacy in bandit problems with partial information.
problem Privacy constraints in multi-armed bandit problems with partial reward information.
method Proposed a generic framework for designing ε ε ε -global DP extensions of UCB and KL-UCB algorithms. result AdaP-KLUCB algorithm achieves optimal regret bound under ε ε ε -global DP constraints. New method reduces multi-armed bandit regret to near-optimal levels.
problem Improving regret bounds for KL-regularized multi-armed bandits.
method Sharp analysis of KL-UCB with peeling argument.
result First high-probability regret bound with linear dependence on K.
This study analyzes mutual influence on investment strategies of financial market agents.
problem Mutual influence among agents in financial markets and its impact on investment strategies.
method Formulated optimal investment differential game problem, derived analytical solutions, proposed fast algorithm, and theoretically analyzed mutual influence.
result Agents' optimal strategies converge to the asymptotic strategy when mutual influence is strong and approaches infinity.
Trading strategies are limited by position limits, leading to a finite number of unique strategies.
problem Limiting the number of long and short positions in trading strategies.
method Formulas and distributions derived for the number of unique trading strategies, transactions, and do-nothing actions.
result A discrete distribution of actions and their properties are presented.
Paper proposes a new framework for combining investment strategies without market-specific assumptions.
problem Lack of a distribution-free and consistent preference framework for decision-making in combining investment strategies.
method Introduces a novel framework for decision-making in combining strategies, free from market conditions and statistical assumptions.
result Proposed strategies outperform individual component strategies in long-term wealth accumulation, with small tradeoffs in Sharpe ratios.
Investment strategies ensure wealth bounded away from zero in a competitive market.
problem Ensuring wealth bounded away from zero in a competitive investment market.
method Stochastic game-theoretic model with survival strategies.
result Survival strategies are asymptotically equivalent and allow faster wealth accumulation.
Study improves StarCraft bot's strategy selection with partial observations.
problem Selecting effective strategies in real-time strategy games with limited information.
method Utilized full game state information during training to predict opponent strategies.
result Substantial win rate improvements over a fixed-strategy baseline.
This paper introduces strategies to maximize arbitrage profits in decentralized exchanges.
problem Maximizing profits from arbitrage loops in decentralized exchanges.
method Three strategies: MaxPrice, MaxMax, and Convex Optimization.
result The Convex Optimization strategy yields the highest monetized arbitrage profit in theory and practice.
In this paper we propose an investing strategy based on neural network models combined with ideas from game-theoretic probability of Shafer and Vovk. Our proposed strategy uses parameter values of a neural network with the best performance until the previous round (trading day) for deciding the investment in the curren…
Automated strategies improve model adaptation efficiency.
problem Manual adaptation strategies are time-consuming and costly.
method Flexible adaptive mechanism deployment for automated adaptation strategies.
result Automated strategies achieve better or comparable performance.
New trading strategies yield gains on average in various market scenarios.
problem Developing trading strategies that consistently yield positive gains in different market conditions.
method Introducing generalized statistical arbitrage concepts and profitable strategies based on information systems.
result Constructed profitable generalized strategies with good performance on simulated and real market data.
Stratify unifies and improves multi-step forecasting strategies.
problem Lack of unified frameworks for multi-step forecasting strategies.
method Proposes Stratify, a parameterized framework for multi-step forecasting.
result Novel strategies in Stratify outperform existing ones in over 84% of experiments.
New method selects best exploration strategies in uncertain environments.
problem Selecting optimal strategies in unknown, multi-strategy environments.
method Formulates Multi-Armed Bandits problem with diversity of effects as reward signal.
result Method outperforms fixed mixtures of strategies in diverse, challenging conditions.
Study finds mean reversion strategies perform well on historical data but fail in recent market conditions.
problem Performance of mean reversion strategies in recent market data.
method Empirical investigation of three mean reversion strategies (PAMR, OLMAR, TCO) on historical S&P 500 data and benchmark datasets.
result Mean reversion strategies may fail in recent market conditions, especially with transaction costs.
Study examines volatility-based strategy for Chinese ETF options, improving returns in volatile markets.
problem Lack of effective trading strategies in volatile Chinese equity markets.
method Volatility forecasting using GARCH models to dynamically adjust positions and exposures.
result Dynamic adjustment of positions and exposures enhances returns in volatile markets.
Model shows how heterogeneity in strategies and risk tolerance affects financial market stability.
problem Understanding how heterogeneity impacts financial market dynamics.
method Agent-based model incorporating heterogeneous investment strategies and risk tolerance.
result Heterogeneity in strategies and risk tolerance suppresses price fluctuations.
Optimal order execution strategies for brokers under reference benchmarks.
problem Maximizing broker's utility of excess profit-and-loss subject to reference strategies.
method Formulated as a utility maximization problem, optimal strategies derived in closed form.
result General reference strategies can be approximated by piece-wise linear combinations of IS and TC orders.
The paper develops a mathematical model for strategic shifts.
problem Finding optimal moments for strategy changes in market dynamics.
method Explicit strategy formulation using fluctuation theory.
result Analytical results predict optimal strategy shifts.
New trading strategy beats traditional grid in crypto markets.
problem Low expected return of traditional grid trading strategy.
method Dynamic Grid Trading (DGT) strategy that adapts to market conditions.
result DGT strategy outperforms traditional grid and buy-and-hold strategies.
A new approach to continuous-time universal portfolios using pathwise Itô calculus.
problem Continuous-time version of Cover's universal portfolio strategies.
method Pathwise Itô calculus approach to establish existence and properties of universal portfolio strategies.
result The universal portfolio strategy's portfolio value process is the average of all values of constant rebalanced strategies.
New method solves continuous time mean-variance model for consistent investment strategy.
problem Time-consistent optimal strategy for continuous time mean-variance model.
method Developed a new Bellman principle method.
result Obtained a time-consistent dynamic optimal strategy.
Paper introduces dynamic strategies for multi-period investment models.
problem Optimizing investment strategies over multiple periods with risk and return considerations.
method Developed a Bellman principle for discrete time multi-period mean-variance models, leading to dynamic optimal strategies and efficient frontiers.
result Dynamic optimal strategies can achieve higher returns with lower risk compared to the 1/n strategy.
We introduce a new general framework for constructing the best trading strategy for a given historical indicator. We construct the unique trading strategy with the highest expected return. This optimal strategy may be implemented directly, or its expected return may be used as a benchmark to evaluate how far away from …
A strategy ensures maximal wealth growth in competitive asset markets.
problem Maximizing wealth growth in competitive asset markets.
method Game-theoretic model and proof of existence of a submartingale strategy.
result Existence and uniqueness of a submartingale strategy that maximizes wealth growth.
A game theory study on optimal hiding and searching strategies in discrete locations.
problem Optimal hiding and searching strategies in a two-person zero-sum game between a hider and a searcher.
method Proved the existence of optimal strategies, developed an algorithm to compute them, and compared with a simple strategy.
result Optimal hiding strategy involves hiding in each location with nonzero probability, and optimal searching strategy can be constructed with up to n simple sequences.
Global optimization in Bayesian inference yields little additional benefit.
problem Improving psychometric parameter estimation using global optimization strategies.
method Experimental simulations comparing myopic and global strategies in multiple models.
result Global optimization strategies provide negligible additional utility improvement beyond the immediate next steps.
The author proposes a finance trading strategy named Entropy Oriented Trading and apply thermodynamics on the strategy. The state variables are chosen so that the strategy satisfies the second law of thermodynamics. Using the law, the author proves that the rate of investment (ROI) of the strategy is equal to or more t…
Study optimal growth strategies in a continuous-time asset market.
problem Guaranteeing that individual agent strategies cannot outperform the market.
method Mean-field approximation of an infinite number of infinitesimal agents, focusing on optimal strategy distribution among assets.
result Optimal strategy for market agents is to invest proportionally to discounted expected relative dividend intensities.
Investment strategy optimized under wealth limits for exponential utility maximization.
problem Maximizing wealth under fixed upper and lower limits for exponential utility.
method Combining optimal investment strategy with options to handle constraints.
result Investment strategy distribution analyzed for change of quantiles.
Survival strategies in a market with self-determined prices are closely tied to log-optimal investment.
problem Survival of wealth in a market with endogenous prices.
method Assume only one's actions affect prices, use log-optimal strategy, disregard actual prices.
result Survival strategies are asymptotically close to log-optimal strategies.
Investigates optimal portfolio strategies in markets with latent side information.
problem Investment problem in markets with latent dependence structure and side information.
method Dynamic and constant portfolio strategies, analyzing log-optimal portfolio as benchmark.
result Optimal dynamic strategy growth rate asymptotically converges to constant strategy in stationary markets.
The aim of this paper is to compare the performances of the optimal strategy under parameters mis-specification and of a technical analysis trading strategy. The setting we consider is that of a stochastic asset price model where the trend follows an unobservable Ornstein-Uhlenbeck process. For both strategies, we prov…