Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

118235353470 · Jun 202019922001200920182026
48 results for minimal exploration

New algorithm reduces regret by allowing free exploration in multi-armed bandits.

problem Designing an adaptive policy to minimize regret with a free exploration budget.
method Introduced (α,β)(α,β)-probably saving policies and a two-phase algorithm UFE-KLUCB-H.
result UFE-KLUCB-H accumulates strictly less regret than non-free exploration policies.

Algorithm achieves optimal pricing with minimal exploration for dynamic markets.

problem Optimal pricing in dynamic markets with contextual information.
method Localized exploration-then-commit (LetC) algorithm with pure exploration, refinement, and exploitation stages.
result Achieves minimax optimal, dimension-free regret bound.

The paper explores stable surfaces in Einstein-Maxwell theory, proving mass bounds and nonexistence results.

problem Exploring stable surfaces in static Einstein-Maxwell space-time.
method Using mean-stable surfaces theory to prove properties of lapse functions and mass bounds.
result Proves ADM mass is bounded by Hawking quasi-local mass.

The study examines minimal surfaces in Riemannian products of surfaces.

problem Exploring geometric and topological restrictions on minimal surfaces in Riemannian products of surfaces.
method Analyzes totally geodesic surfaces and minimal 2-spheres, 2-tori, and 2-spheres in Riemannian products of surfaces with constant curvature.
result Generically, a totally geodesic surface in a Riemannian product is either a slice or a product of geodesics. Minimal 2-spheres and 2-tori have specific properties under certain curvature conditions.

In graph-based active learning, algorithms based on expected error minimization (EEM) have been popular and yield good empirical performance. The exact computation of EEM optimally balances exploration and exploitation. In practice, however, EEM-based algorithms employ various approximations due to the computational ha…

2016-09-03abs ↗pdf ↗

This paper explores neural network loss landscapes and their effects on generalization.

problem Understanding the structure of neural network loss functions and their impact on generalization.
method Simple filter normalization and various visualization methods to explore loss landscape structure and network architecture effects.
result Visualizations reveal how network architecture and training parameters affect loss landscape curvature and minimizers.

A new method combines online and offline learning to tackle contextual bandits with missing action support.

problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.

We explore a connection between the Finslerian area functional and well-investigated Cartan functionals to prove new Bernstein theorems, uniqueness and removability results for Finsler-minimal graphs, as well as enclosure theorems and isoperimetric inequalities for minimal immersions in Finsler spaces. In addition, we …

2014-03-31abs ↗pdf ↗

New algorithms improve exploration in MDPs with theoretical guarantees.

problem Efficient exploration in undiscounted MDPs with continuous states.
method Exploration bonuses for SCAL and C-SCAL algorithms.
result Achieves sublinear regret with improved computational efficiency.

Study reward-free RL in non-linear settings, improving efficiency and removing assumptions.

problem Improving sample efficiency in reward-free reinforcement learning for non-linear function approximation.
method Proposed RFOLIVE algorithm for minimal structural assumptions, analyzed hardness results for reward-free and reward-aware exploration.
result Statistical efficiency and hardness results under various structural assumptions, no need for reachability or explorability assumptions.

The paper explores risk-minimization for exponential additive models, providing mathematical expressions and numerical examples.

problem Risk-minimization in incomplete markets for exponential additive models.
method Derive explicit mathematical expressions for local risk-minimization strategies in exponential additive models.
result Provide necessary conditions for deriving expressions and confirm integrability conditions for specific models.

This paper explores twisted Lagrangian tori in C^2 and their Hamiltonian stationarity.

problem Understanding the Hamiltonian stationarity of twisted Lagrangian tori in C^2.
method Investigation of differential geometry of twisted tori, including product and Chekanov's exotic tori.
result Only product tori are minimal under Hamiltonian deformations, indicating Chekanov's exotic tori are not area minimal.

Minimal hypersurfaces can't always be connected by mean curvature flow.

problem Existence of connecting mean curvature flows for minimal hypersurfaces.
method Minimal hypersurface analogue of gradient flow trajectories between critical points.
result Additional topological and variational obstructions to connecting mean curvature flows.

This paper explores portfolio management strategies to maximize alpha and minimize beta.

problem Maximizing returns while minimizing risk in investment portfolios.
method Examines asset allocation, diversification, active management, and risk management strategies.
result Combining these strategies optimizes portfolio performance.

Proposes a new method for finding frequent closed patterns in transaction bases.

problem Frequent closed patterns in transaction bases.
method Partitioning the search space into subcontexts and updating frequent closed patterns with their minimal generators.
result Proposed approach called UFCIGs-DAC for efficient search of frequent closed itemsets.

Algorithm optimizes two objectives in bandits: minimizing regret and identifying best arm.

problem Balancing exploration and exploitation for optimal performance in multi-armed bandits.
method Design and analysis of BoBW-lil'UCB(γ)(γ) algorithm, establishing lower bounds.
result BoBW-lil'UCB(γ)(γ) achieves optimal performance for RM or BAI under different γγ values.

New findings show increased exploration needed in non-stationary RL tasks.

problem Task non-stationarity leads to conflicting goals in RL.
method Analyzes the trade-off between cumulative and simple regret in non-stationary environments.
result Increased exploration is necessary to balance CR and SR in non-stationary tasks.

Algorithm reduces regret in partially observable systems by learning dynamics and using optimistic control.

problem Minimizing regret in partially observable linear quadratic control systems with unknown dynamics.
method ExpCommit algorithm that learns model parameters and uses optimism in uncertainty.
result End-to-end sublinear regret upper bound of O~(T2/3)\tilde{\mathcal{O}}(T^{2/3}) for ExpCommit.

New algorithm eliminates arms to minimize regret in complex bandit problems.

problem Minimizing regret in combinatorial bandit problems with explicit exploration.
method Introduces a novel arm elimination scheme that partitions arms into three categories and incorporates explicit exploration.
result Achieves near-optimal regret in combinatorial multi-armed and linear contextual bandit problems.

Proposes a method to avoid excessive exploration in reinforcement learning.

problem Avoiding excessive exploration in reinforcement learning to deploy it in practice.
method Designs a novel algorithm using UCB reinforcement learning policy with adaptive exploration constraints.
result Proves that the approach remains conservative while minimizing regret in tabular settings and validates on real-world tasks.

In this article we explore some finer properties of equi-areal mirrors and introduce techniques for developing new mirror surfaces that simultaneously minimize angular and areal distortion.

2014-10-24abs ↗pdf ↗

Automated meta-learning for contextual bandits improves efficiency and performance.

problem Optimizing decision-making in dynamic environments like personalization and recommendation systems.
method End-to-end automated meta-learning pipeline using linearly annealed e-greedy policy.
result Model outperforms random exploration and other models with minimal tuning.

RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for PO with mediator feedback.

problem Policy Optimization in continuous control tasks.
method RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for regret minimization in PO.
result Achieving constant regret under certain circumstances in PO with mediator feedback.

In this paper, we explore minimal contact triangulations on contact 3-manifolds. We give many explicit examples of contact triangulations that are close to minimal ones. The main results of this article say that on any closed oriented 3-manifold the number of vertices for minimal contact triangulations for overtwisted …

2016-08-12abs ↗pdf ↗

The paper minimizes Borda regret in dueling bandits models.

problem Minimizing Borda regret in dueling bandits models.
method Proposes explore-then-commit and EXP3-type algorithms for stochastic and adversarial settings respectively.
result Achieves nearly matching regret upper bounds of O(d2/3T2/3)O(d^{2/3} T^{2/3}) for both settings.

AdaLinUCB optimizes exploration-exploitation for contextually varying costs.

problem Optimizing decision-making in environments with varying exploration costs.
method Adaptive Upper-Confidence-Bound (AdaLinUCB) algorithm for opportunistic learning.
result AdaLinUCB achieves O((log T)^2) regret bound, significantly outperforming other algorithms.

Minimal assumptions analysis of Q-learning with time-varying policies.

problem Finite-time analysis of Q-learning with time-varying policies for discounted MDPs.
method Minimal assumptions, Poisson equation decomposition, sensitivity analysis.
result Established convergence rate and sample complexity for Q-learning.

A new method boosts exploration in bandit algorithms, reducing regret.

problem Improving exploration in bandit algorithms with bounded or unbounded rewards.
method Residual Bootstrap Exploration (ReBoot) method that injects data-driven randomness.
result Proves logarithmic regret in Gaussian multi-armed bandits with appropriate variance inflation.

The paper explores minimal coloring numbers for Z\mathbb{Z}-colorable links.

problem Finding the minimum number of colors needed for Z\mathbb{Z}-colorings on minimal diagrams of Z\mathbb{Z}-colorable links.
method Investigates minimal diagrams and Z\mathbb{Z}-colorings for Z\mathbb{Z}-colorable links.
result For any positive integer NN, there exists a minimal diagram of a Z\mathbb{Z}-colorable link with at least NN colors in any Z\mathbb{Z}-coloring.

We consider the problem of minimizing capital at risk in the Black-Scholes setting. The portfolio problem is studied given the possibility that a correlation constraint between the portfolio and a financial index is imposed. The optimal portfolio is obtained in closed form. The effects of the correlation constraint are…

2014-11-24abs ↗pdf ↗