New algorithm reduces regret by allowing free exploration in multi-armed bandits.
problem Designing an adaptive policy to minimize regret with a free exploration budget.
method Introduced (α,β)-probably saving policies and a two-phase algorithm UFE-KLUCB-H. result UFE-KLUCB-H accumulates strictly less regret than non-free exploration policies.
Algorithm achieves optimal pricing with minimal exploration for dynamic markets.
problem Optimal pricing in dynamic markets with contextual information.
method Localized exploration-then-commit (LetC) algorithm with pure exploration, refinement, and exploitation stages.
result Achieves minimax optimal, dimension-free regret bound.
New algorithm reduces exploration in structured stochastic bandits.
problem Wide class of stochastic bandit problems with known structural properties.
method Developed OSSB algorithm that matches minimal exploration rates of sub-optimal arms.
result OSSB's regret matches asymptotic instance-specific regret lower bound.
The paper explores stable surfaces in Einstein-Maxwell theory, proving mass bounds and nonexistence results.
problem Exploring stable surfaces in static Einstein-Maxwell space-time.
method Using mean-stable surfaces theory to prove properties of lapse functions and mass bounds.
result Proves ADM mass is bounded by Hawking quasi-local mass.
The study examines minimal surfaces in Riemannian products of surfaces.
problem Exploring geometric and topological restrictions on minimal surfaces in Riemannian products of surfaces.
method Analyzes totally geodesic surfaces and minimal 2-spheres, 2-tori, and 2-spheres in Riemannian products of surfaces with constant curvature.
result Generically, a totally geodesic surface in a Riemannian product is either a slice or a product of geodesics. Minimal 2-spheres and 2-tori have specific properties under certain curvature conditions.
In graph-based active learning, algorithms based on expected error minimization (EEM) have been popular and yield good empirical performance. The exact computation of EEM optimally balances exploration and exploitation. In practice, however, EEM-based algorithms employ various approximations due to the computational ha…
This paper explores neural network loss landscapes and their effects on generalization.
problem Understanding the structure of neural network loss functions and their impact on generalization.
method Simple filter normalization and various visualization methods to explore loss landscape structure and network architecture effects.
result Visualizations reveal how network architecture and training parameters affect loss landscape curvature and minimizers.
New algorithm reduces regret in contextual bandits.
problem Minimizing regret in contextual bandits with side information.
method Contextual-Gap algorithm for simple regret minimization.
result Established performance guarantees on simple regret.
Paper explores cohomology classes on non-compact almost Kähler manifolds.
problem Understanding cohomology classes induced by symplectic forms.
method Provides criteria for non-trivial classes in Lp cohomology. result Symplectic forms induce non-trivial classes in Lp cohomology. FLEX optimizes exploration for nonlinear systems with minimal data.
problem Efficient exploration of unknown nonlinear systems with limited data.
method FLEX uses optimal experimental design to maximize information gain.
result FLEX outperforms other methods in nonlinear environments and control tasks.
A new method combines online and offline learning to tackle contextual bandits with missing action support.
problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.
We explore a connection between the Finslerian area functional and well-investigated Cartan functionals to prove new Bernstein theorems, uniqueness and removability results for Finsler-minimal graphs, as well as enclosure theorems and isoperimetric inequalities for minimal immersions in Finsler spaces. In addition, we …
New algorithms improve exploration in MDPs with theoretical guarantees.
problem Efficient exploration in undiscounted MDPs with continuous states.
method Exploration bonuses for SCAL and C-SCAL algorithms.
result Achieves sublinear regret with improved computational efficiency.
ISL algorithm tackles deep exploration efficiently.
problem Deep exploration in reinforcement learning.
method Derives ISL algorithm by augmenting RL objective with a novel regularization term.
result Empirically shows state-of-the-art performance on deep-exploration benchmarks.
Study reward-free RL in non-linear settings, improving efficiency and removing assumptions.
problem Improving sample efficiency in reward-free reinforcement learning for non-linear function approximation.
method Proposed RFOLIVE algorithm for minimal structural assumptions, analyzed hardness results for reward-free and reward-aware exploration.
result Statistical efficiency and hardness results under various structural assumptions, no need for reachability or explorability assumptions.
The paper explores risk-minimization for exponential additive models, providing mathematical expressions and numerical examples.
problem Risk-minimization in incomplete markets for exponential additive models.
method Derive explicit mathematical expressions for local risk-minimization strategies in exponential additive models.
result Provide necessary conditions for deriving expressions and confirm integrability conditions for specific models.
Framework for interactive learning to minimize user experience regressions.
problem Sub-optimal user experiences due to frequent exploration of options.
method Explore-Exploit framework for online learning operators.
result Efficiencies achieved in integrating online learning with run-time services.
This paper explores twisted Lagrangian tori in C^2 and their Hamiltonian stationarity.
problem Understanding the Hamiltonian stationarity of twisted Lagrangian tori in C^2.
method Investigation of differential geometry of twisted tori, including product and Chekanov's exotic tori.
result Only product tori are minimal under Hamiltonian deformations, indicating Chekanov's exotic tori are not area minimal.
Paper tackles robust control policy learning for uncertain systems.
problem Learning control policies for an unknown linear dynamical system with quadratic cost.
method Convex optimization method balancing exploitation and exploration.
result Minimizes worst-case cost by reducing uncertainty in model parameters.
Study calculates spectra of minimal hypersurfaces in hyperbolic space.
problem Computing Laplacian spectra of minimal hypersurfaces.
method Analyzes hypersurfaces in hyperbolic space with specific asymptotic data.
result Obtains spectra and extremal properties of the bottom of the spectrum.
Minimal hypersurfaces can't always be connected by mean curvature flow.
problem Existence of connecting mean curvature flows for minimal hypersurfaces.
method Minimal hypersurface analogue of gradient flow trajectories between critical points.
result Additional topological and variational obstructions to connecting mean curvature flows.
This paper explores portfolio management strategies to maximize alpha and minimize beta.
problem Maximizing returns while minimizing risk in investment portfolios.
method Examines asset allocation, diversification, active management, and risk management strategies.
result Combining these strategies optimizes portfolio performance.
Proposes a new method for finding frequent closed patterns in transaction bases.
problem Frequent closed patterns in transaction bases.
method Partitioning the search space into subcontexts and updating frequent closed patterns with their minimal generators.
result Proposed approach called UFCIGs-DAC for efficient search of frequent closed itemsets.
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
problem Escaping local minima in loss landscapes.
method Hill-ADAM alternates between minimizing and maximizing error to explore the loss space.
result Hill-ADAM finds the global minimum state in loss landscapes.
Algorithm optimizes two objectives in bandits: minimizing regret and identifying best arm.
problem Balancing exploration and exploitation for optimal performance in multi-armed bandits.
method Design and analysis of BoBW-lil'UCB(γ) algorithm, establishing lower bounds. result BoBW-lil'UCB(γ) achieves optimal performance for RM or BAI under different γ values. Minimal hyperbolic surface diameter grows logarithmically with genus.
problem Finding the smallest possible diameter of hyperbolic surfaces.
method Random construction, lattice point counting, and exploration of random trivalent graphs.
result Minimal diameter is asymptotic to log(g) as genus g approaches infinity.
New findings show increased exploration needed in non-stationary RL tasks.
problem Task non-stationarity leads to conflicting goals in RL.
method Analyzes the trade-off between cumulative and simple regret in non-stationary environments.
result Increased exploration is necessary to balance CR and SR in non-stationary tasks.
Improved Thompson Sampling for Bayesian Optimization.
problem Handling the exploitation-exploration dilemma in Bayesian optimization.
method Incorporating epsilon-greedy policy into Thompson Sampling.
result Epsilon-greedy Thompson Sampling outperforms standard TS extremes.
Algorithm reduces regret in partially observable systems by learning dynamics and using optimistic control.
problem Minimizing regret in partially observable linear quadratic control systems with unknown dynamics.
method ExpCommit algorithm that learns model parameters and uses optimism in uncertainty.
result End-to-end sublinear regret upper bound of O~(T2/3) for ExpCommit. New algorithm eliminates arms to minimize regret in complex bandit problems.
problem Minimizing regret in combinatorial bandit problems with explicit exploration.
method Introduces a novel arm elimination scheme that partitions arms into three categories and incorporates explicit exploration.
result Achieves near-optimal regret in combinatorial multi-armed and linear contextual bandit problems.
Proposes a method to avoid excessive exploration in reinforcement learning.
problem Avoiding excessive exploration in reinforcement learning to deploy it in practice.
method Designs a novel algorithm using UCB reinforcement learning policy with adaptive exploration constraints.
result Proves that the approach remains conservative while minimizing regret in tabular settings and validates on real-world tasks.
In this article we explore some finer properties of equi-areal mirrors and introduce techniques for developing new mirror surfaces that simultaneously minimize angular and areal distortion.
Algorithm reduces online regret by leveraging offline data in linear bandits.
problem Online regret minimization in linear bandits with offline data.
method OOPE algorithm using extended D-optimal design.
result Substantial reduction in online regret compared to prior work.
Automated meta-learning for contextual bandits improves efficiency and performance.
problem Optimizing decision-making in dynamic environments like personalization and recommendation systems.
method End-to-end automated meta-learning pipeline using linearly annealed e-greedy policy.
result Model outperforms random exploration and other models with minimal tuning.
Algorithm reduces exploration in structured RL problems.
problem Minimize exploration in reinforcement learning with known structure.
method Directed Exploration Learning (DEL) for Lipschitz MDPs.
result Regret lower bounds not scaling with state and action space sizes.
RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for PO with mediator feedback.
problem Policy Optimization in continuous control tasks.
method RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for regret minimization in PO.
result Achieving constant regret under certain circumstances in PO with mediator feedback.
A new method learns to stop with minimal data, outperforming traditional approaches.
problem Optimal stopping problems with unknown distributions.
method Explore-then-exploit approach with logarithmic exploration phase.
result Performance comparable to full information DP solution with minimal exploration.
Recent progress on minimal surface system and cones in Euclidean spaces.
problem Exploring the Dirichlet problem for minimal surfaces and cones.
method Systematic developments and new families of minimizing cones.
result New families of minimizing cones of different types.
A new algorithm reduces regret in multiplayer bandits with minimal communication.
problem Maximizing rewards in multiplayer multi-armed bandits with collisions.
method DPE (Decentralized Parsimonious Exploration) algorithm.
result Achieves the same regret as optimal centralized algorithms with less communication.
In this paper, we explore minimal contact triangulations on contact 3-manifolds. We give many explicit examples of contact triangulations that are close to minimal ones. The main results of this article say that on any closed oriented 3-manifold the number of vertices for minimal contact triangulations for overtwisted …
The paper minimizes Borda regret in dueling bandits models.
problem Minimizing Borda regret in dueling bandits models.
method Proposes explore-then-commit and EXP3-type algorithms for stochastic and adversarial settings respectively.
result Achieves nearly matching regret upper bounds of O(d2/3T2/3) for both settings. AdaLinUCB optimizes exploration-exploitation for contextually varying costs.
problem Optimizing decision-making in environments with varying exploration costs.
method Adaptive Upper-Confidence-Bound (AdaLinUCB) algorithm for opportunistic learning.
result AdaLinUCB achieves O((log T)^2) regret bound, significantly outperforming other algorithms.
Minimal assumptions analysis of Q-learning with time-varying policies.
problem Finite-time analysis of Q-learning with time-varying policies for discounted MDPs.
method Minimal assumptions, Poisson equation decomposition, sensitivity analysis.
result Established convergence rate and sample complexity for Q-learning.
A new method boosts exploration in bandit algorithms, reducing regret.
problem Improving exploration in bandit algorithms with bounded or unbounded rewards.
method Residual Bootstrap Exploration (ReBoot) method that injects data-driven randomness.
result Proves logarithmic regret in Gaussian multi-armed bandits with appropriate variance inflation.
E4 algorithm optimizes batched linear bandits with minimal regret and batches.
problem Optimizing batched linear bandits for minimal regret.
method Explore-Estimate-Eliminate-Exploit framework with optimal exploration rate.
result Achieves minimax and asymptotic optimality in regret and batch complexity.
MIME uses mutual information minimization for better exploration in environments with abrupt transitions.
problem Agents struggle at abrupt environmental transitions.
method MIME learns a latent representation without predicting future states.
result MIME outperforms surprisal-driven agents at transition boundaries.
The paper explores minimal coloring numbers for Z-colorable links.
problem Finding the minimum number of colors needed for Z-colorings on minimal diagrams of Z-colorable links. method Investigates minimal diagrams and Z-colorings for Z-colorable links. result For any positive integer N, there exists a minimal diagram of a Z-colorable link with at least N colors in any Z-coloring. We consider the problem of minimizing capital at risk in the Black-Scholes setting. The portfolio problem is studied given the possibility that a correlation constraint between the portfolio and a financial index is imposed. The optimal portfolio is obtained in closed form. The effects of the correlation constraint are…