Thompson sampling, a Bayesian method for balancing exploration and exploitation in bandit problems, has theoretical guarantees and exhibits strong empirical performance in many domains. Traditional Thompson sampling, however, assumes perfect compliance, where an agent's chosen action is treated as the implemented actio…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian Causal Forest models estimate treatment effects with noncompliance.
New method improves treatment effect estimation in adaptive experiments with noncompliance.
BRACE addresses noncompliance in bandits, offering methods for recommendation and treatment policies.
Motivated by clinical trials, we study bandits with observable non-compliance. At each step, the learner chooses an arm, after, instead of observing only the reward, it also observes the action that took place. We show that such noncompliance can be helpful or hurtful to the learner in general. Unfortunately, naively i…
A microscopic dynamic model is here constructed and analyzed, describing the evolution of the income distribution in the presence of taxation and redistribution in a society in which also tax evasion and auditing processes occur. The focus is on effects of enforcement regimes, characterized by different choices of the …
The paper addresses causal mediation analysis with post-treatment events, proposing robust estimators and efficient methods.
We extend the classic multi-armed bandit (MAB) model to the setting of noncompliance, where the arm pull is a mere instrument and the treatment applied may differ from it, which gives rise to the instrument-armed bandit (IAB) problem. The IAB setting is relevant whenever the experimental units are human since free will…
There is a fast-growing literature on estimating optimal treatment regimes based on randomized trials or observational studies under a key identifying condition of no unmeasured confounding. Because confounding by unmeasured factors cannot generally be ruled out with certainty in observational studies or randomized tri…
Automated method bounds causal effects in discrete data.
New algorithm learns optimal policies in strategic MDPs with private types.
Develops framework for estimating and improving DTRs with time-varying IV in the presence of unmeasured confounding.