New algorithm for contextual combinatorial bandits with probabilistic arm triggering.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We analyze the regret of combinatorial Thompson sampling (CTS) for the combinatorial multi-armed bandit with probabilistically triggered arms under the semi-bandit feedback setting. We assume that the learner has access to an exact optimization oracle but does not know the expected base arm outcomes beforehand. When th…
Paper improves CMAB regret bounds by reducing batch-size dependency.
We study combinatorial multi-armed bandit with probabilistically triggered arms (CMAB-T) and semi-bandit feedback. We resolve a serious issue in the prior CMAB-T studies where the regret bounds contain a possibly exponentially large factor of , where is the minimum positive probability that an arm is trigg…
In this paper, we study the combinatorial multi-armed bandit problem (CMAB) with probabilistically triggered arms (PTAs). Under the assumption that the arm triggering probabilities (ATPs) are positive for all arms, we prove that a class of upper confidence bound (UCB) policies, named Combinatorial UCB with exploration …
Influence maximization, adaptive routing, and dynamic spectrum allocation all require choosing the right action from a large set of alternatives. Thanks to the advances in combinatorial optimization, these and many similar problems can be efficiently solved given an environment with known stochasticity. In this paper, …
Graph-Triggered Bandits unify rested and restless bandits with graph-defined arm interactions.
Study improves CTS's approximation regret for combinatorial bandits.
ET-GP-UCB optimizes time-varying functions without knowing change rates.
New algorithm minimizes regret in multi-agent bandit problem with probabilistic communication.
We consider a problem of stochastic online learning with general probabilistic graph feedback, where each directed edge in the feedback graph has probability . Two cases are covered. (a) The one-step case, where after playing arm the learner observes a sample reward feedback of arm with independent prob…
New algorithm for multi-player bandits with selfish players, achieving logarithmic regret.
New algorithm tackles non-stationary combinatorial semi-bandit problems with optimal regret bounds.
Combines offline and online learning for identifying the best arm in bandits.
We present a probabilistic model of events in continuous time in which each event triggers a Poisson process of successor events. The ensemble of observed events is thereby modeled as a superposition of Poisson processes. Efficient inference is feasible under this model with an EM algorithm. Moreover, the EM algorithm …
New Thompson Sampling for partially observed context bandits reduces regret logarithmically with time.
Motivated by models of human decision making proposed to explain commonly observed deviations from conventional expected value preferences, we formulate two stochastic multi-armed bandit problems with distorted probabilities on the reward distributions: the classic -armed bandit and the linearly parameterized bandit…
MINTS uses a minimalist Bayesian framework to tackle multi-armed bandits with structural constraints.
A new Bayesian framework simplifies stochastic optimization by focusing on key parameters.
Study best arm identification with contextual info, achieving optimal misidentification probability.
Study on reward poisoning attacks on CMAB, revealing attackability depends on adversary's knowledge.
Paper improves voice trigger detection for privacy-centric smart assistants.
Neural backdoor attack is emerging as a severe security threat to deep learning, while the capability of existing defense methods is limited, especially for complex backdoor triggers. In the work, we explore the space formed by the pixel values of all possible backdoor triggers. An original trigger used by an attacker …
Pricing Chinese convertible bonds using Monte Carlo simulation and dynamic programming.
Paper solves time-inconsistent control problems with BSDEs.
MISA detects Trojan triggers in neural networks at inference time.
Common event-triggered state estimation (ETSE) algorithms save communication in networked control systems by predicting agents' behavior, and transmitting updates only when the predictions deviate significantly. The effectiveness in reducing communication thus heavily depends on the quality of the dynamics models used …
New graph feedback model for bandits with improved regret bounds.
PHAZE framework uses zkML and hashing for fast, verifiable LHC trigger decisions.
Improved voice trigger detection in noisy environments.
Unified framework for distributional regret in bandits and reinforcement learning.
We consider the combinatorial multi-armed bandit (CMAB) problem, where the reward function is nonlinear. In this setting, the agent chooses a batch of arms on each round and receives feedback from each arm of the batch. The reward that the agent aims to maximize is a function of the selected arms and their expectations…
CascadeBAI identifies best arms in cascading bandits with fixed confidence.
PPPD framework extracts physical characterizations from stochastic mechanical systems.
Thompson Sampling improves decision-making in partially observed contexts.
Backdoor attacks make models predict a specific class near triggers, smoothing their decision function.
LinConTS improves regret and constraint violations in probabilistic linearly constrained bandits.
We deliver a call to arms for probabilistic numerical methods: algorithms for numerical tasks, including linear algebra, integration, optimization and solving differential equations, that return uncertainties in their calculations. Such uncertainties, arising from the loss of precision induced by numerical calculation …
We introduce a dynamic mechanism for the solution of analytically-tractable substructure in probabilistic programs, using conjugate priors and affine transformations to reduce variance in Monte Carlo estimators. For inference with Sequential Monte Carlo, this automatically yields improvements such as locally-optimal pr…
The need for consistent treatment of uncertainty has recently triggered increased interest in probabilistic deep learning methods. However, most current approaches have severe limitations when it comes to inference, since many of these models do not even permit to evaluate exact data likelihoods. Sum-product networks (…
We consider the problem of building a state representation model for control, in a continual learning setting. As the environment changes, the aim is to efficiently compress the sensory state's information without losing past knowledge, and then use Reinforcement Learning on the resulting features for efficient policy …
BadGD identifies gradient descent vulnerabilities through strategic backdoor attacks.
TS-Insight visualizes Thompson Sampling for better debugging and trust.
Corporate defaults may be triggered by some major market news or events such as financial crises or collapses of major banks or financial institutions. With a view to develop a more realistic model for credit risk analysis, we introduce a new type of reduced-form intensity-based model that can incorporate the impacts o…
Improved speech recognition for voice assistants by analyzing speech data.
In classical Hawkes process, the baseline intensity and triggering kernel are assumed to be a constant and parametric function respectively, which limits the model flexibility. To generalize it, we present a fully Bayesian nonparametric model, namely Gaussian process modulated Hawkes process and propose an EM-variation…
We describe a methodology for modeling the performance of decision-level data fusion between different sensor configurations, implemented as part of the JIEDDO Analytic Decision Engine (JADE). We first discuss a Bayesian network formulation of classical probabilistic data fusion, which allows elementary fusion structur…
We design a new myopic strategy for a wide class of sequential design of experiment (DOE) problems, where the goal is to collect data in order to to fulfil a certain problem specific goal. Our approach, Myopic Posterior Sampling (MPS), is inspired by the classical posterior (Thompson) sampling algorithm for multi-armed…