Comparison Lift uses bandit algorithms to optimize online ad testing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces active and passive causal inference techniques.
Modern deep learning methods are very sensitive to many hyperparameters, and, due to the long training times of state-of-the-art models, vanilla Bayesian hyperparameter optimization is typically computationally infeasible. On the other hand, bandit-based configuration evaluation approaches based on random search lack g…
This paper investigates the grant-free random access with massive IoT devices. By embedding the data symbols in the signature sequences, joint device activity detection and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms that enjoy a low…
AdamCB optimizes neural network training by adaptively selecting samples.
Paper proposes an efficient bandit-based algorithm for hyperparameter optimization.
The celebrated Monte Carlo method estimates an expensive-to-compute quantity by random sampling. Bandit-based Monte Carlo optimization is a general technique for computing the minimum of many such expensive-to-compute quantities by adaptive random sampling. The technique converts an optimization problem into a statisti…
Bayesian optimization improved for biased data.
Performance of machine learning algorithms depends critically on identifying a good set of hyperparameters. While recent approaches use Bayesian optimization to adaptively select configurations, we focus on speeding up random search through adaptive resource allocation and early-stopping. We formulate hyperparameter op…
We propose a contextual bandit based model to capture the learning and social welfare goals of a web platform in the presence of myopic users. By using payments to incentivize these agents to explore different items/recommendations, we show how the platform can learn the inherent attributes of items and achieve a subli…
We consider an online decision making setting known as contextual bandit problem, and propose an approach for improving contextual bandit performance by using an adaptive feature extraction (representation learning) based on online clustering. Our approach starts with an off-line pre-training on unlabeled history of co…
Improved UCB algorithm for diversity in bandits with lower bounds.
Given a huge set of applicants, how should a firm allocate sequential resume screenings, phone interviews, and in-person site visits? In a tiered interview process, later stages (e.g., in-person visits) are more informative, but also more expensive than earlier stages (e.g., resume screenings). Using accepted hiring mo…
New algorithm tunes SGMCMC hyperparameters for scalable Bayesian inference.
The paper presents a method for personalized exercise recommendations that improves learner skill gain.
BLiE optimizes hyperparameters with theoretical guarantees and superior performance.
A new bandit algorithm for web page item display.
We study the problem of black-box optimization of a noisy function in the presence of low-cost approximations or fidelities, which is motivated by problems like hyper-parameter tuning. In hyper-parameter tuning evaluating the black-box function at a point involves training a learning algorithm on a large data-set at a …
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these applications is critical for protecting online user experiences but very challenging due to their "p…
Automated algorithm selection and hyperparameter tuning facilitates the application of machine learning. Traditional multi-armed bandit strategies look to the history of observed rewards to identify the most promising arms for optimizing expected total reward in the long run. When considering limited time budgets and c…
New framework uses tree ensembles for contextual bandits.
DPPS uses DP priors for Bayesian non-parametric multi-arm bandits.
Database activity monitoring (DAM) systems are commonly used by organizations to protect the organizational data, knowledge and intellectual properties. In order to protect organizations database DAM systems have two main roles, monitoring (documenting activity) and alerting to anomalous activity. Due to high-velocity …
Next generation networks are expected to be ultradense and aim to explore spectrum sharing paradigm that allows users to communicate in licensed, shared as well as unlicensed spectrum. Such ultra-dense networks will incur significant signaling load at base stations leading to a negative effect on spectrum and energy ef…
A contextual bandit method evaluates and improves inventory control policies.
ProtoBandit uses bandits to find prototypes efficiently.
A blockchain protocol uses bandit algorithms to dynamically price transactions.
Two algorithms improve GP bandits by selecting priors and minimizing regret.
Precision oncology, the genetic sequencing of tumors to identify druggable targets, has emerged as the standard of care in the treatment of many cancers. Nonetheless, due to the pace of therapy development and variability in patient information, designing effective protocols for individual treatment assignment in a sam…
Unified approach for conversational recommendation by integrating attributes and items.
New UCB-type algorithms reduce regret bounds for stochastic bandits with heavy and super heavy noise.
Gamblers lose in long bets despite casino claims, study shows.
Optimum-statistical collaboration improves black-box optimization efficiency.
A new algorithm for personalized recommendations adapts to changing user interests.
Pioneers a new network slicing solution for beyond-5G networks.
Autonomous cyber-physical agents and systems play an increasingly large role in our lives. To ensure that agents behave in ways aligned with the values of the societies in which they operate, we must develop techniques that allow these agents to not only maximize their reward in an environment, but also to learn and fo…
Improved UCB method for stochastic bandits using distance tuning.
BLOB combines organic and bandit signals for better user interest estimation.
Adaptive speculative decoding framework for LLMs using bandit algorithms.
Unified framework for scalable black-box optimization.
MBExplainer provides explanations for models combining graph embeddings and tabular features.
Adaptive algorithm selects models for stochastic linear bandits based on problem complexity.
Geometric approach combines asset returns and investor views for better portfolio optimization.
Paper proposes an alternative method to price American options using HJM approach.
We study inference and learning based on a sparse coding model with `spike-and-slab' prior. As in standard sparse coding, the model used assumes independent latent sources that linearly combine to generate data points. However, instead of using a standard sparse prior such as a Laplace distribution, we study the applic…
Quantum machine learning: Adiabatic quantum SVM outperforms classical methods.
We develop a semi-analytic approach to the valuation of auto-callable structures with accrual features subject to barrier conditions. Our approach is based on recent studies of multi-assed binaries, present in the literature. We extend these studies to the case of time-dependent parameters. We compare numerically the s…
Two ML approaches learn local volatility surfaces from option prices, with GP being arbitrage-free.