Paper develops a dynamic Bayesian approach for active learning that optimizes exploration-exploitation balance.
problem Balancing exploration and exploitation in active learning for unknown functions.
method Develops BHEEM, a Bayesian hierarchical approach with approximate Bayesian computation for sampling trade-off parameters.
result BHEEM achieves at least 21% and 11% improvement over pure exploration and exploitation strategies respectively.
Batch Thompson Sampling reduces exploration-exploitation trade-off in online decision making.
problem Balancing exploration and exploitation in online decision making.
method Introducing a batch Thompson Sampling framework for stochastic multi-arm bandit and linear contextual bandit problems.
result Achieves asymptotic regret bound with O ( log T ) O(\log T) O ( log T ) batch queries, significantly reducing interactions. Survey on methods for balancing exploration and exploitation in reinforcement learning.
problem Balancing exploration and exploitation in reinforcement learning, especially in domains with limited data.
method Survey of methods for computing robust solutions from fixed samples.
result Presentation of methods for balancing exploration-exploitation trade-off.
Bayesian bandits use double sampling to balance exploration and exploitation.
problem Balancing exploration and exploitation in real-world systems.
method Develops a double sampling technique to balance exploration and exploitation in Bayesian settings.
result Empirically shows reduced cumulative regret compared to state-of-the-art alternatives.
NEXT learns efficient paths in high dimensions using neural exploration-exploitation trees.
problem Learning efficient path planning in high-dimensional spaces.
method Neural Exploration-Exploitation Trees (NEXT) integrating neural architecture and UCB algorithm.
result NEXT achieves better sample efficiency and outperforms state-of-the-art methods.
Active learning method improves AI performance by balancing exploration and exploitation.
problem Efficiently acquiring samples for supervised learning in streaming data.
method Ensemble active learning by contextual bandits.
result Improved AI modeling performance through better sample acquisition.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
problem Exploration-exploitation dilemma in finite-horizon reinforcement learning with metric state-action spaces.
method Kernel-UCBVI, leveraging smoothness and kernel estimators of rewards and transitions.
result First regret bound for kernel-based RL using smoothing kernels, O ( H 3 K 2 d / ( 2 d + 1 ) ) O(H^3 K^{2d/(2d+1)}) O ( H 3 K 2 d / ( 2 d + 1 ) ) . Meta-SAC automatically tunes SAC's entropy temperature for better exploration.
problem Exploration-exploitation dilemma in reinforcement learning.
method Meta-SAC uses metagradient and a novel meta objective to automatically adjust SAC's entropy temperature.
result Meta-SAC outperforms SAC-v2 by 10% on the humanoid-v2 task.
This paper applies Thompson Sampling to asymmetric α \alpha α -stable bandits for financial and wireless data.
problem Optimizing exploration-exploitation in multi-armed bandits with asymmetric α \alpha α -stable distributions. method Thompson Sampling applied to unknown asymmetric α \alpha α -stable reward distributions. result Demonstrates effectiveness of Thompson Sampling for asymmetric α \alpha α -stable bandits. Improved control approach for correlated bandits with better performance.
problem General multi-armed bandit problem with correlated elements.
method Introducing entropy regularisation to obtain a smooth asymptotic approximation of the value function, leading to a semi-index approximation of the optimal decision process.
result Performance of Asymptotic Randomised Control (ARC) algorithm compares favorably with other approaches.
AdaLinUCB optimizes exploration-exploitation for contextually varying costs.
problem Optimizing decision-making in environments with varying exploration costs.
method Adaptive Upper-Confidence-Bound (AdaLinUCB) algorithm for opportunistic learning.
result AdaLinUCB achieves O((log T)^2) regret bound, significantly outperforming other algorithms.
This research tackles balancing exploration and exploitation in deep RL for partially observable systems.
problem Balancing exploration and exploitation in deep RL for partially observable systems.
method Deployed and tested several techniques including adaptive and deterministic exploration strategies, and a modified quadratic loss function.
result Adaptive methods better approximate the trade-off between exploration and exploitation.
Graph-based active learning improves with a new algorithm that balances exploration and exploitation.
problem Graph-based active learning algorithms based on expected error minimization (EEM) often use approximations due to computational hardness, leading to suboptimal performance.
method Proposes TSA (Two-Step Approximation) algorithm that efficiently balances exploration and exploitation with similar computational complexity.
result Empirically shows that balancing exploration and exploitation improves performance in both toy and real-world datasets.
New algorithm balances exploration and exploitation in opportunistic bandits.
problem Regret of pulling suboptimal arms varies with environmental conditions.
method Proposes AdaUCB algorithm to adaptively balance exploration and exploitation.
result AdaUCB achieves O ( log T ) O(\log T) O ( log T ) regret with a smaller coefficient than traditional UCB. Multi-armed bandit problems are the most basic examples of sequential decision problems with an exploration-exploitation trade-off. This is the balance between staying with the option that gave highest payoffs in the past and exploring new options that might give higher payoffs in the future. Although the study of band…
Active inference enhances RL by balancing exploration and exploitation.
problem Traditional RL's balance between exploration and exploitation is often suboptimal.
method Developed a new decision-making objective based on active inference.
result The new algorithm successfully balances exploration and exploitation on various RL benchmarks.
The paper analyzes CMDPs, balancing exploration and exploitation to avoid constraint violations.
problem Balancing exploration and exploitation in CMDPs to satisfy constraints.
method Two approaches: optimistic planning and incremental updates of primal and dual variables.
result Both approaches achieve sublinear regret on utility and constraint violations, with stronger guarantees for the linear programming approach.
Algorithm achieves optimal pricing with minimal exploration for dynamic markets.
problem Optimal pricing in dynamic markets with contextual information.
method Localized exploration-then-commit (LetC) algorithm with pure exploration, refinement, and exploitation stages.
result Achieves minimax optimal, dimension-free regret bound.
Algorithm balances online and offline data for linear bandits.
problem Online learning with an offline dataset in linear bandits.
method Proposes a linear bandit algorithm that uses offline data early and increasingly favors exploration as the horizon grows.
result Establishes regret bounds showing competitive performance with both purely online and offline solutions.
Info-p strategy optimally balances exploration and exploitation in slot machines.
problem Balancing exploration and exploitation in slot machines.
method Infomax strategy (Info-p) to maximize information.
result Info-p strategy optimally saturates known optimal bounds and compares favorably to existing policies.
New method learns adaptive exploration strategies for dynamic tasks.
problem Learning effective exploration strategies in changing environments.
method Informed policy regularization to reduce sample complexity of RNN-based policies.
result Method learns efficient exploration strategies balancing information gathering and reward maximization.
LaMBO optimizes biological sequences using autoencoders and Bayesian optimization.
problem Bayesian optimization for drug design is limited by discrete, high-dimensional decision variables.
method Jointly trains denoising autoencoder with a Gaussian process head for gradient-based optimization in latent space.
result LaMBO outperforms genetic optimizers and requires no large pretraining corpus.
AgABC improves ABC algorithm by balancing exploration and exploitation.
problem Balancing global and local search abilities in ABC algorithm.
method Divide population into groups and assign different search strategies to members.
result Proposed AgABC algorithm outperforms other algorithms in accuracy and stability.
Improved analysis of UCRL2 with empirical Bernstein inequality reduces exploration-exploitation regret.
problem Exploration-exploitation in communicating Markov Decision Processes.
method Analysis of UCRL2 with Empirical Bernstein inequalities (UCRL2B).
result Regret bound of O ~ ( D Γ S A T ) \widetilde{O}(\sqrt{DΓS A T}) O ( D Γ S A T ) for UCRL2B. AIS algorithm balances exploration and exploitation for efficient sampling.
problem Balancing exploration and exploitation in adaptive importance sampling.
method Daisee algorithm, partition-based approach, pseudo-regret analysis.
result Daisee achieves O ( T ( log T ) 3 4 ) \mathcal{O}(\sqrt{T}(\log T)^{\frac{3}{4}}) O ( T ( log T ) 4 3 ) cumulative pseudo-regret. DiffATD efficiently discovers targets in partially observable environments using diffusion dynamics.
problem Efficiently discovering targets in partially observable environments with limited sampling.
method DiffATD uses diffusion dynamics to maintain a belief distribution over unobserved states, balancing exploration and exploitation.
result DiffATD outperforms baselines and supervised methods in diverse domains.
Proposes a new algorithm for efficient online model selection of LLMs considering the increasing-then-converging trend.
problem Balancing cost and performance in choosing the best LLM among a diverse set of models.
method Introduces a time-increasing bandit algorithm (TI-UCB) that predicts model performance increases and balances exploration and exploitation.
result Achieves a logarithmic regret upper bound, indicating a fast convergence rate in model selection.
A new algorithm for cryo-EM data collection that balances reward and latency.
problem Optimizing data collection in cryo-EM experiments with action delays.
method Latency-aware contextual bandit framework and COAF algorithm.
result The COAF algorithm efficiently maximizes reward over time in cryo-EM experiments.
New AIM algorithm optimizes exploration-exploitation in bandits.
problem Balancing exploration and exploitation in decision-making.
method Approximate Information Maximization (AIM) algorithm.
result AIM outperforms Infomax and Thompson sampling with enhanced speed and tractability.
FPQL improves Q-learning for quantum system control.
problem Balancing exploration and exploitation in quantum control.
method Fidelity-based probabilistic Q-learning (FPQL) for quantum systems.
result FPQL achieves better balance and avoids local optima.
TS-Insight visualizes Thompson Sampling for better debugging and trust.
problem Thompson Sampling's black box nature hinders debugging and trust.
method TS-Insight is a visual analytics tool that traces evolving posteriors and evidence counts.
result Visualizations help in verifying, diagnosing, and explaining Thompson Sampling dynamics.
Budgeted hyper-parameter tuning algorithm improves performance.
problem Optimizing hyper-parameters with resource constraints.
method Sequential decision making, Bayesian model, action-value function.
result Superior performance across various budgets.
An agent explores indefinitely in an environment with unlimited rewards.
problem Balancing exploration and exploitation in environments with unlimited rewards.
method Simple example of an environment with unbounded rewards and optimal agent behavior.
result An optimal agent always explores to maximize rewards, regardless of accumulated knowledge.
Deep Bayesian Bandits compare methods for balancing exploration and exploitation in complex domains.
problem Balancing exploration and exploitation in complex sequential decision-making problems.
method Benchmarking approximate Bayesian neural networks with Thompson Sampling over contextual bandit problems.
result Many approaches successful in supervised learning underperform in sequential decision-making.
New algorithm learns POMDPs with known observation model efficiently.
problem Learning POMDPs with unknown transition model in average-reward setting.
method OAS estimation technique and OAS-UCRL algorithm balancing exploration-exploitation.
result Regret guarantee of order O ( T log ( T ) ) \mathcal{O}(\sqrt{T \log(T)}) O ( T log ( T ) ) for OAS-UCRL algorithm. Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.
We develop a coherent framework for integrative simultaneous analysis of the exploration-exploitation and model order selection trade-offs. We improve over our preceding results on the same subject (Seldin et al., 2011) by combining PAC-Bayesian analysis with Bernstein-type inequality for martingales. Such a combinatio…
ICEE learns new RL tasks in less time with a Transformer model.
problem Efficient in-context policy learning for reinforcement learning.
method In-context Exploration-Exploitation (ICEE) algorithm that optimizes efficiency without explicit Bayesian inference.
result ICEE solves Bayesian optimization problems as efficiently as Gaussian process biased methods but in significantly less time.
New algorithm prevents strategic replication in multi-armed bandit problems.
problem Strategic replication by agents can exploit bandit algorithms' balance.
method Designs Hierarchical UCB (H-UCB) and Robust Hierarchical UCB (RH-UCB) algorithms.
result Achieves O ( ln T ) O(\ln T) O ( ln T ) -regret and sublinear regret in realistic scenarios. SCAL algorithm reduces exploration-exploitation in unknown MDPs with bias span constraints.
problem Efficient exploration-exploitation in unknown weakly-communicating MDPs with bias span constraints.
method Introduces SCAL, an algorithm that proves a regret bound of O(c√(ΓSAT)) for unknown MDPs with known bias span.
result SCAL significantly outperforms existing algorithms like UCRL and PSRL in MDPs with large diameter and small bias span.
New methods optimize wireless base station tuning in dense networks.
problem Optimizing wireless base station parameters in dense cellular networks.
method Parallel contextual bandit approaches using Thompson sampling and deterministic UCB.
result Thompson sampling outperforms manual tuning and contextual UCB in real base station network data.
New strategies improve multi-agent decision-making on irregular networks.
problem Maximizing group reward in multi-agent settings with heterogeneous strategies.
method Design and analysis of heterogeneous explore-exploit strategies for multi-star networks.
result Group performance improves under heterogeneous strategies compared to homogeneous strategies.
Maximize to Explore integrates RL components for efficient policy discovery.
problem Balancing exploration and exploitation in online RL with general function approximators.
method Integrates estimation, planning, and exploration into a single objective function.
result Achieves sublinear regret for MDPs and MGs with general function approximations.
A method for high-dimensional Bayesian optimization reduces dimensionality using EDR and Gaussian process.
problem Extending Bayesian optimization to high-dimensional settings.
method Two-step framework: EDR subspace identification followed by Gaussian process optimization.
result Algorithm converges in high-dimensional contexts, validated by numerical experiments.
New algorithms balance collaboration and adversarial behavior in linear bandits.
problem Minimizing regret in a collaborative linear bandit problem with adversarial agents.
method Robust collaborative phased elimination algorithm with tight analyses.
result Achieves near-optimal regret bounds of $O\left(α+ 1/\sqrt{M}
ight) \sqrt{dT}$ for good agents.
Safe-Optimization algorithm shows participants' preference for safe outcomes over optimal points.
problem Exploration-exploitation in functions with safety constraints.
method Safe-Optimization algorithm based on Gaussian Processes, tested in two experiments.
result Participants prioritize safety over optimization, showing a homeostatic strategy.
Study examines impact of missing data on multi-armed bandit algorithms.
problem Impact of missing data on performance of multi-armed bandit algorithms.
method Extensive simulation study of two-armed bandit algorithms with binary outcomes, considering different probabilities of missingness.
result Impact on performance varies depending on the balance between exploration and exploitation.
Study explores strategies for randomized allocation in delayed rewards bandits.
problem Understanding the exploration-exploitation tradeoff in randomized strategies with delayed rewards.
method Examines two strategies: updating exploration sequence at every time point vs. updating only when a new reward is observed.
result The strategy updating only when a new reward is observed leads to strong consistency in allocation for a wider scope of situations.