Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

96191287382 · Jun 202019922001200920172026
48 results for finite exploration

Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.

problem Exploration-exploitation dilemma in finite-horizon reinforcement learning with metric state-action spaces.
method Kernel-UCBVI, leveraging smoothness and kernel estimators of rewards and transitions.
result First regret bound for kernel-based RL using smoothing kernels, O(H3K2d/(2d+1))O(H^3 K^{2d/(2d+1)}).

In this paper, we study multi-armed bandit problems in explore-then-commit setting. In our proposed explore-then-commit setting, the goal is to identify the best arm after a pure experimentation (exploration) phase and exploit it once or for a given finite number of times. We identify that although the arm with the hig…

2019-04-30abs ↗pdf ↗

Safe actions learned in finite trials, without infinite exploration.

problem Learning safe actions in unknown environments efficiently.
method Defining a handicap metric and using sequential probability ratio test for discarding unsafe actions.
result Achieves constant handicap, discarding unsafe machines with probability one in finite rounds.

Study on endomorphism and automorphism groups of specific quandles.

problem Characterizing endomorphism and automorphism groups of residually finite and profinite quandles.
method Proved properties of endomorphism monoids and automorphism groups for residually finite and profinite quandles.
result Endomorphism and automorphism groups of residually finite quandles are residually finite.

We study the exploration problem in episodic MDPs with rich observations generated from a small number of latent states. Under certain identifiability assumptions, we demonstrate how to estimate a mapping from the observations to latent states inductively through a sequence of regression and clustering steps -- where p…

2019-01-25abs ↗pdf ↗

We present a 1-parameter family of finite action solutions to the S0(2,1)S0(2,1) Hitchin's equations and explore some of its basic properties. For a fixed value of the parameter, the solution is smooth. We conclude by showing a multi-particle generalization of our basic solutions.

2000-11-13abs ↗pdf ↗

RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for PO with mediator feedback.

problem Policy Optimization in continuous control tasks.
method RANDomized-exploration policy Optimization via Multiple Importance Sampling with Truncation (RANDOMIST) for regret minimization in PO.
result Achieving constant regret under certain circumstances in PO with mediator feedback.

We examine groups whose resonance varieties, characteristic varieties and Sigma-invariants have a natural arithmetic group symmetry, and we explore implications on various finiteness properties of subgroups. We compute resonance varieties, characteristic varieties and Alexander polynomials of Torelli groups, and we sho…

2010-02-03abs ↗pdf ↗

We consider the class of curves of finite total curvature, as introduced by Milnor. This is a natural class for variational problems and geometric knot theory, and since it includes both smooth and polygonal curves, its study shows us connections between discrete and differential geometry. To explore these ideas, we co…

2006-06-01abs ↗pdf ↗

Finite-dimensional spaces of biharmonic functions on manifolds are explored.

problem Characterizing biharmonic functions on open manifolds with nonnegative Ricci curvature.
method Analyzing bounded and polynomial growth biharmonic functions, deriving Weyl bounds, and studying fourth-order operators.
result Finite-dimensional spaces of biharmonic functions with polynomial growth are established.

The Runge-Kutta-Legendre scheme improves pricing American options and other derivatives.

problem Pricing American options and other derivatives with improved accuracy and stability.
method Runge-Kutta-Legendre finite difference scheme applied to Black-Scholes and Heston models.
result Improved convergence and stability compared to existing schemes.

Ghost points affect stability in finite difference schemes for diffusion equations.

problem Impact of ghost points on stability of finite difference schemes.
method Exploration of explicit Euler finite difference scheme with ghost points on diffusion equation.
result Stability of the scheme is affected by ghost points.

This paper approaches the definition and properties of dynamic convex risk measures through the notion of a family of concave valuation operators satisfying certain simple and credible axioms. Exploring these in the simplest context of a finite time set and finite sample space, we find natural risk-transfer and time-co…

2007-09-03abs ↗pdf ↗

Contextual bandits have the same exploration-exploitation trade-off as standard multi-armed bandits. On adding positive externalities that decay with time, this problem becomes much more difficult as wrong decisions at the start are hard to recover from. We explore existing policies in this setting and highlight their …

2019-11-14abs ↗pdf ↗

Study explores strategies for randomized allocation in delayed rewards bandits.

problem Understanding the exploration-exploitation tradeoff in randomized strategies with delayed rewards.
method Examines two strategies: updating exploration sequence at every time point vs. updating only when a new reward is observed.
result The strategy updating only when a new reward is observed leads to strong consistency in allocation for a wider scope of situations.

Characterizes monodromies of projective structures on finite-type surfaces.

problem Understanding monodromies of projective structures on finite-type surfaces.
method Geometrical/topological study of local conical projective structures.
result Any representation can be represented as the holonomy of a branched projective structure.

In this paper, we explore holomorphic Segre preserving maps. First, we investigate holomorphic Segre preserving maps sending the complexification M\mathcal{M} of a generic real analytic submanifold $M \subseteq \C^N$ of finite type at some point pp into the complexification M\mathcal{M}' of a generic real analytic s…

2008-10-14abs ↗pdf ↗

Study finds finitely many non-congruent polygonal domains with same Steklov spectrum.

problem Inverse Steklov problem on convex polygons.
method Analysis of Steklov eigenvalues and isoperimetric bounds.
result For almost all convex polygonal domains, there exist at most finitely many non-congruent domains with the same Steklov spectrum.

Gradient Ricci solitons can be extended to non-gradient Ricci solitons using energy function.

problem Extending the geometry of gradient Ricci solitons to non-gradient Ricci solitons.
method Using energy function EE to study the geometry.
result A non-steady Ricci soliton with symmetric covariant derivative is gradient.

Minimal assumptions analysis of Q-learning with time-varying policies.

problem Finite-time analysis of Q-learning with time-varying policies for discounted MDPs.
method Minimal assumptions, Poisson equation decomposition, sensitivity analysis.
result Established convergence rate and sample complexity for Q-learning.

New approach incentivizes strategic agents to explore, making exploration almost free.

problem Incentivized exploration in multi-armed bandits with long-term strategic agents.
method Simple incentive-provision strategy, best arm identification algorithm, and UCB lower bound.
result Exploration can be (almost) free when there are many learning agents.

Study task-guided exploration in linear dynamical systems, improving sample complexity.

problem Efficiently learning about an environment to complete a specific task.
method Proposed a computationally efficient experiment-design based exploration algorithm.
result Optimally explores the environment, collecting precise information needed to complete the task.

Study of deep neural networks using finite-time Lyapunov exponents.

problem Understanding the geometric structures in input space formed by deep neural networks.
method Analogy with dynamical systems, computing finite-time Lyapunov exponents.
result Ridges of large positive exponents divide input space into regions associated with different classes.

The paper explores geometric finiteness in mapping class groups and constructs new examples of these subgroups.

problem Understanding geometric finiteness in mapping class groups and constructing new examples.
method Examined several constructions of subgroups and determined conditions for geometric finiteness.
result Provides new examples of parabolically geometrically finite and reducibly geometrically finite subgroups.

The study explores finite quotients of 3-manifold groups and their existence and non-existence.

problem Does there exist a 3-manifold group with a specific finite quotient but not others?
method The approach combines group cohomology, topological results, and probabilistic methods.
result Proves existence and non-existence of 3-manifolds with certain finite quotients.

Control of non-episodic, finite-horizon dynamical systems with uncertain dynamics poses a tough and elementary case of the exploration-exploitation trade-off. Bayesian reinforcement learning, reasoning about the effect of actions and future observations, offers a principled solution, but is intractable. We review, then…

2015-10-13abs ↗pdf ↗

This paper studies the problem of identifying any kk distinct arms among the top ρρ fraction (e.g., top 5\%) of arms from a finite or infinite set with a probably approximately correct (PAC) tolerance εε. We consider two cases: (i) when the threshold of the top arms' expected rewards is known and (ii) when it is unk…

2018-10-28abs ↗pdf ↗

Paper explores fair classification with bounded disparity using finite datasets.

problem Ensuring fairness in binary classification with protected groups.
method Minimax optimal approach with fairness constraints and demographic disparity control.
result Proposes FairBayes-DDP+ method that achieves minimax lower bound on fairness-aware excess risk.

DE is a new exploration method that limits resource usage based on expected improvement and surprise.

problem Limited exploration in large action spaces when resources are scarce.
method Delight-gated exploration (DE) that limits exploration actions based on a gate price set by the product of expected improvement and surprise.
result DE outperforms ε\varepsilon-greedy and Thompson Sampling in terms of regret across various bandit and MDP settings.

The study of infinite groups through their finite quotients in geometry.

problem Understanding properties of infinite groups from their finite images.
method Analyzing infinite groups through their finite quotients and using low-dimensional topology.
result Recent results show how finite images can determine the group completely in some cases.