Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4693139185 · Jun 202019922001200920182026
48 results for KL-UCB strategy

KL-UCB-switch optimizes bandit strategies for both distribution-dependent and distribution-free performance.

problem Optimizing regret bounds for stochastic bandits.
method Combining MOSS and KL-UCB strategies.
result Achieves both optimal distribution-dependent and distribution-free regret bounds.

New algorithm outperforms existing ones in multi-player bandit problems without sensing.

problem Decentralized multi-player multi-armed bandit problem without collision or sensing info.
method Randomized Selfish KL-UCB, inspired by Selfish KL-UCB, with low complexity.
result Randomized Selfish KL-UCB outperforms state-of-the-art algorithms in almost all environments.

New findings show many popular bandit algorithms are unstable, contradicting minimax optimality.

problem Challenges in statistical inference from bandit algorithms due to adaptive, non-i.i.d. nature.
method Analysis of stability properties of optimism-based bandit algorithms.
result Widely used minimax-optimal UCB-style algorithms are unstable.

This paper is about index policies for minimizing (frequentist) regret in a stochastic multi-armed bandit model, inspired by a Bayesian view on the problem. Our main contribution is to prove that the Bayes-UCB algorithm, which relies on quantiles of posterior distributions, is asymptotically optimal when the reward dis…

2016-01-06abs ↗pdf ↗

The profitable bandit problem aims to maximize earnings by choosing actions with uncertain rewards.

problem Maximizing earnings from uncertain rewards in a subset selection problem.
method Adapted and studied three strategies: kl-UCB, Bayes-UCB, and Thompson Sampling.
result Establishes asymptotic optimality for each strategy with finite time regret bounds.

We propose the kl-UCB ++ algorithm for regret minimization in stochastic bandit models with exponential families of distributions. We prove that it is simultaneously asymptotically optimal (in the sense of Lai and Robbins' lower bound) and minimax optimal. This is the first algorithm proved to enjoy these two propertie…

2017-02-23abs ↗pdf ↗

We study a generalization of the multi-armed bandit problem with multiple plays where there is a cost associated with pulling each arm and the agent has a budget at each time that dictates how much she can expect to spend. We derive an asymptotic regret lower bound for any uniformly efficient algorithm in our setting. …

2016-06-30abs ↗pdf ↗

The paper analyzes the sliding regret of stochastic bandit algorithms.

problem Measuring the one-shot behavior of no-regret algorithms in stochastic bandits.
method Introducing sliding regret to measure the worst pseudo-regret over a time-window.
result Randomized methods have optimal sliding regret, while index policies have the worst possible sliding regret.

A new algorithm for better decision-making in recommendation systems.

problem Stochastic multi-armed bandit problem and cold start problem in recommender systems.
method Proposes Hellinger-UCB, a variant of UCB algorithm using squared Hellinger distance.
result Hellinger-UCB reaches the theoretical lower bound and outperforms other algorithms in practical applications.

New algorithms improve privacy in bandit problems with partial information.

problem Privacy constraints in multi-armed bandit problems with partial reward information.
method Proposed a generic framework for designing εε-global DP extensions of UCB and KL-UCB algorithms.
result AdaP-KLUCB algorithm achieves optimal regret bound under εε-global DP constraints.

This study analyzes mutual influence on investment strategies of financial market agents.

problem Mutual influence among agents in financial markets and its impact on investment strategies.
method Formulated optimal investment differential game problem, derived analytical solutions, proposed fast algorithm, and theoretically analyzed mutual influence.
result Agents' optimal strategies converge to the asymptotic strategy when mutual influence is strong and approaches infinity.

Trading strategies are limited by position limits, leading to a finite number of unique strategies.

problem Limiting the number of long and short positions in trading strategies.
method Formulas and distributions derived for the number of unique trading strategies, transactions, and do-nothing actions.
result A discrete distribution of actions and their properties are presented.

Paper proposes a new framework for combining investment strategies without market-specific assumptions.

problem Lack of a distribution-free and consistent preference framework for decision-making in combining investment strategies.
method Introduces a novel framework for decision-making in combining strategies, free from market conditions and statistical assumptions.
result Proposed strategies outperform individual component strategies in long-term wealth accumulation, with small tradeoffs in Sharpe ratios.

Investment strategies ensure wealth bounded away from zero in a competitive market.

problem Ensuring wealth bounded away from zero in a competitive investment market.
method Stochastic game-theoretic model with survival strategies.
result Survival strategies are asymptotically equivalent and allow faster wealth accumulation.

Study improves StarCraft bot's strategy selection with partial observations.

problem Selecting effective strategies in real-time strategy games with limited information.
method Utilized full game state information during training to predict opponent strategies.
result Substantial win rate improvements over a fixed-strategy baseline.

In this paper we propose an investing strategy based on neural network models combined with ideas from game-theoretic probability of Shafer and Vovk. Our proposed strategy uses parameter values of a neural network with the best performance until the previous round (trading day) for deciding the investment in the curren…

2010-02-11abs ↗pdf ↗

New trading strategies yield gains on average in various market scenarios.

problem Developing trading strategies that consistently yield positive gains in different market conditions.
method Introducing generalized statistical arbitrage concepts and profitable strategies based on information systems.
result Constructed profitable generalized strategies with good performance on simulated and real market data.

New method selects best exploration strategies in uncertain environments.

problem Selecting optimal strategies in unknown, multi-strategy environments.
method Formulates Multi-Armed Bandits problem with diversity of effects as reward signal.
result Method outperforms fixed mixtures of strategies in diverse, challenging conditions.

Study finds mean reversion strategies perform well on historical data but fail in recent market conditions.

problem Performance of mean reversion strategies in recent market data.
method Empirical investigation of three mean reversion strategies (PAMR, OLMAR, TCO) on historical S&P 500 data and benchmark datasets.
result Mean reversion strategies may fail in recent market conditions, especially with transaction costs.

Study examines volatility-based strategy for Chinese ETF options, improving returns in volatile markets.

problem Lack of effective trading strategies in volatile Chinese equity markets.
method Volatility forecasting using GARCH models to dynamically adjust positions and exposures.
result Dynamic adjustment of positions and exposures enhances returns in volatile markets.

Model shows how heterogeneity in strategies and risk tolerance affects financial market stability.

problem Understanding how heterogeneity impacts financial market dynamics.
method Agent-based model incorporating heterogeneous investment strategies and risk tolerance.
result Heterogeneity in strategies and risk tolerance suppresses price fluctuations.

Optimal order execution strategies for brokers under reference benchmarks.

problem Maximizing broker's utility of excess profit-and-loss subject to reference strategies.
method Formulated as a utility maximization problem, optimal strategies derived in closed form.
result General reference strategies can be approximated by piece-wise linear combinations of IS and TC orders.

A new approach to continuous-time universal portfolios using pathwise Itô calculus.

problem Continuous-time version of Cover's universal portfolio strategies.
method Pathwise Itô calculus approach to establish existence and properties of universal portfolio strategies.
result The universal portfolio strategy's portfolio value process is the average of all values of constant rebalanced strategies.

Paper introduces dynamic strategies for multi-period investment models.

problem Optimizing investment strategies over multiple periods with risk and return considerations.
method Developed a Bellman principle for discrete time multi-period mean-variance models, leading to dynamic optimal strategies and efficient frontiers.
result Dynamic optimal strategies can achieve higher returns with lower risk compared to the 1/n strategy.

We introduce a new general framework for constructing the best trading strategy for a given historical indicator. We construct the unique trading strategy with the highest expected return. This optimal strategy may be implemented directly, or its expected return may be used as a benchmark to evaluate how far away from …

2011-08-03abs ↗pdf ↗

A game theory study on optimal hiding and searching strategies in discrete locations.

problem Optimal hiding and searching strategies in a two-person zero-sum game between a hider and a searcher.
method Proved the existence of optimal strategies, developed an algorithm to compute them, and compared with a simple strategy.
result Optimal hiding strategy involves hiding in each location with nonzero probability, and optimal searching strategy can be constructed with up to n simple sequences.

Global optimization in Bayesian inference yields little additional benefit.

problem Improving psychometric parameter estimation using global optimization strategies.
method Experimental simulations comparing myopic and global strategies in multiple models.
result Global optimization strategies provide negligible additional utility improvement beyond the immediate next steps.

Study optimal growth strategies in a continuous-time asset market.

problem Guaranteeing that individual agent strategies cannot outperform the market.
method Mean-field approximation of an infinite number of infinitesimal agents, focusing on optimal strategy distribution among assets.
result Optimal strategy for market agents is to invest proportionally to discounted expected relative dividend intensities.

Investment strategy optimized under wealth limits for exponential utility maximization.

problem Maximizing wealth under fixed upper and lower limits for exponential utility.
method Combining optimal investment strategy with options to handle constraints.
result Investment strategy distribution analyzed for change of quantiles.

Survival strategies in a market with self-determined prices are closely tied to log-optimal investment.

problem Survival of wealth in a market with endogenous prices.
method Assume only one's actions affect prices, use log-optimal strategy, disregard actual prices.
result Survival strategies are asymptotically close to log-optimal strategies.

Investigates optimal portfolio strategies in markets with latent side information.

problem Investment problem in markets with latent dependence structure and side information.
method Dynamic and constant portfolio strategies, analyzing log-optimal portfolio as benchmark.
result Optimal dynamic strategy growth rate asymptotically converges to constant strategy in stationary markets.

The aim of this paper is to compare the performances of the optimal strategy under parameters mis-specification and of a technical analysis trading strategy. The setting we consider is that of a stochastic asset price model where the trend follows an unobservable Ornstein-Uhlenbeck process. For both strategies, we prov…

2016-04-30abs ↗pdf ↗