Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1223 · Mar 202019922001200920182026
21 results for exploitation-exploration

New method combines population and completion tasks in knowledge graphs.

problem Insufficient external resources hinder statistical inference in knowledge graphs.
method Probabilistic factorisation method that uses path structure for both population and completion.
result Balanced exploitation-exploration helps incremental population and improves prediction of missing information.

XploVAE improves recommendation by balancing known and novel items.

problem Balancing known and novel items for better recommendations.
method Constructs user-specific subgraphs for exploitation and exploration, learns personalized item embeddings.
result Demonstrates effectiveness on various real-world datasets.

Paper explores active learning strategies for real-time credit card fraud detection.

problem Challenges in labeling and imbalanced transaction data for real-time fraud detection.
method Investigates active learning strategies for querying unlabeled transactions, comparing supervised, semi-supervised, and unsupervised approaches.
result Highlights an exploitation/exploration trade-off for active learning in fraud detection.

Proposes EE-Net for neural exploration in contextual bandits.

problem Exploitation-Exploration tradeoff in contextual bandits.
method Uses two neural networks: Exploitation and Exploration, to learn reward function and adaptively explore.
result Achieves O(TlogT)\mathcal{O}(\sqrt{T\log T}) regret and outperforms existing methods.

Pseudo-Bayesian Optimization improves black-box function optimization using simple local regression.

problem Optimizing expensive black-box functions with uncertainty quantification.
method Axiomatic framework for convergence guarantees, combined with simple local regression and randomized prior.
result Empirically superior algorithms outperform state-of-the-art benchmarks in various applications.

ZoomRL learns efficient strategies for large state-action spaces using a metric.

problem Handling large state-action spaces in reinforcement learning.
method ZoomRL leverages continuous bandits to adaptively discretize the joint space.
result Achieves worst-case regret of $ ilde{O}(H^{ rac{5}{2}} K^{ rac{d+1}{d+2}})$.

New algorithm optimizes online decision-making with dynamically generated actions.

problem Balancing action generation costs with optimal decision-making in online learning.
method Doubly-optimistic algorithm using LCB for action selection and UCB for action generation.
result Achieves optimal regret bound of O(Tdd+2ddd+2+dTlogT)O(T^{\frac{d}{d+2}}d^{\frac{d}{d+2}} + d\sqrt{T\log T}).

AGG-UCB uses neural networks to optimize group behaviors in contextual bandits.

problem Optimizing group behaviors in contextual bandits with mutual impacts.
method Introduces Arm Group Graph (AGG) and AGG-UCB algorithm using neural networks and graph neural networks.
result Achieves near-optimal regret bound with over-parameterized neural networks.

LetSIP learns relevant patterns for user interests in data mining.

problem Redundancy in pattern mining makes it hard for analysts to identify relevant patterns.
method Combines pattern sampling with interactive data mining, using user feedback to learn sampling distribution.
result Favourable trade-offs in quality-diversity and exploitation-exploration compared to existing methods.

DiffATD efficiently discovers targets in partially observable environments using diffusion dynamics.

problem Efficiently discovering targets in partially observable environments with limited sampling.
method DiffATD uses diffusion dynamics to maintain a belief distribution over unobserved states, balancing exploration and exploitation.
result DiffATD outperforms baselines and supervised methods in diverse domains.

Conversational UCB accelerates bandit learning with user feedback.

problem Slow learning speed in traditional contextual bandit algorithms.
method Generalized contextual bandit to conversational contextual bandit, leveraging both behavioral and conversational feedbacks.
result ConUCB achieves a smaller regret upper bound, indicating faster learning speed.

MatrixRL tackles RL in high-dimensional spaces with feature and kernel methods, achieving near-optimal regret bounds.

problem Challenges of exploration in RL with large state-action spaces.
method MatrixRL combines linear bandit techniques with feature and kernel methods to learn a low-dimensional representation of the transition model.
result MatrixRL achieves an O(H2dlogTT){O}\big(H^2d\log T\sqrt{T}\big) regret bound, near-optimal in time TT and dimension dd.

Aims to optimize influence spread in social networks using bandit algorithms.

problem Maximizing influence spread in unknown social networks.
method Combines Thompson Sampling and Epsilon Greedy algorithms with automatic ensemble learning.
result Demonstrates effectiveness of automatic ensemble learning for combinatorial bandit problems.

New algorithm optimizes reward while ensuring safety in complex decision-making problems.

problem Maximizing reward while adhering to safety constraints in complex decision-making problems.
method Optimistic Primal-Dual Proximal Policy Optimization (OPDOP) algorithm combining least-squares policy evaluation and a bonus term for safe exploration.
result Achieves ildeO(dH2.5T) ilde{O}(d H^{2.5}\sqrt{T}) regret and ildeO(dH2.5T) ilde{O}(d H^{2.5}\sqrt{T}) constraint violation.