Proposes stabilized weights for causal inference using isotonic calibration.
problem Stability and bias issues in inverse propensity weighting.
method Post-hoc isotonic calibration of inverse propensity weights.
result Improves performance of doubly robust estimators of average treatment effect.
GEAR uses auxiliary data to estimate optimal decisions in studies with limited primary outcomes.
problem Estimating optimal decisions when primary outcomes are not available in experimental samples.
method GEAR uses augmented inverse propensity weighting to estimate optimal decisions based on auxiliary data.
result GEAR estimators and value estimators have established asymptotic properties and are validated in simulations and a real application.
New framework quantifies error in causal inference methods.
problem Uncertainty in estimating causal effects due to confounders.
method Representation learning approach to characterize error bounds.
result General error bounds for causal inference methods, including IPW.
New method improves confidence intervals for adaptive experiment results.
problem Adaptive experiments complicate statistical inference, especially when estimating sub-optimal treatments.
method Adaptive reweighting of inverse propensity weighting terms to control variance and ensure correct coverage.
result The method prevents heavy-tailed estimates and increases hypothesis testing power.
Paper tackles informative labels in semi-supervised learning, proposing debiasing methods.
problem Informative labels can bias semi-supervised learning models, especially when some classes are more likely to be labeled.
method Estimates missing-data mechanism and uses inverse propensity weighting to debias SSL algorithms.
result Proposed methods improve SSL performance, demonstrated on various datasets including medical ones.
New MLIPS method reduces error in batch bandit learning.
problem Learning from logged bandit feedback with policy discrepancy.
method Estimates a surrogate policy to reduce error in inverse propensity weighting.
result MLIPS has smaller mean squared error than IPS.
Improved off-policy selection and learning in contextual bandits with better guarantees.
problem Selecting or training a reward-maximizing policy using data from a fixed behavior policy.
method A betting-based confidence bound applied to an inverse propensity weight sequence for off-policy selection, and a freezing condition for off-policy learning.
result The proposed methods achieve significantly improved guarantees over prior work, especially in small-data regimes.
DTS improves robustness of bandit algorithms in nonstationary environments.
problem Brittle behavior of multi-armed bandit algorithms in nonstationary exogenous factors.
method Deconfounded Thompson Sampling (DTS) that projects population-level performance while controlling for context.
result DTS provides resilience to exogenous variation and balances exploration and exploitation.
Study designs for estimating treatment effects in adaptive experiments.
problem Estimating treatment effects under adaptive treatment assignment.
method Propose and analyze IPW and AIPW estimators, establish CLTs under design stability.
result Central limit theorems for IPW and AIPW estimators under design stability.
FIDDLE uses deep learning to estimate ATE from complex data.
problem Estimating ATE from high-dimensional, correlated covariates with sparse nonlinear effects.
method Factor-augmented deep learning for propensity and outcome models.
result FIDDLE consistently estimates ATE under model misspecification and is semiparametrically efficient.
Paper tackles exposure bias in recommender systems using contrastive learning.
problem Exposure bias in large-scale recommender systems.
method Contrastive learning to reduce exposure bias via inverse propensity weighting.
result Contrastive learning effectively reduces exposure bias in recommender systems.
Kernelized bandit algorithm tackles adaptive contextual bandits with single-index models.
problem Adaptive contextual bandits with single-index models and unknown link functions.
method Kernelized ε-greedy algorithm combining Stein-based index estimation and kernel ridge regression for reward functions.
result Unified framework for simultaneous learning and inference in single-index contextual bandits.
Adapting policy learning for data collected from evolving systems.
problem Challenges in learning optimal policies from adaptively collected data.
method Proposes an algorithm based on generalized augmented inverse propensity weighted (AIPW) estimators to control worst-case estimation variance.
result Achieves minimax rate optimal regret guarantees even with diminishing exploration.
Procedure for unbiasedly estimating value of optimized policies.
problem Unbiased estimation of value of optimized policies in A/B testing.
method Bagging process with inverse-propensity-weighting and per-sample value estimates.
result Unbiased estimator of the value of deploying an optimized policy.
Proposes rounding method for precise treatment effect estimation under budget constraints.
problem Resource-constrained experimental design for precise treatment effect estimation.
method Dependent randomized rounding procedure to convert assignment probabilities into binary treatment decisions.
result Improved estimator precision through variance reduction and efficient inference.
New method learns decisions from collective preferences without individual covariates.
problem Making decisions online without individual covariates.
method Collaborative filtering, matrix completion bandit, ε-greedy policy, online gradient descent, inverse propensity weighting.
result Method outperforms benchmarks and reveals new discoveries.
R-Learning uses inverse-variance weights to estimate treatment effects more accurately.
problem Estimating heterogeneous treatment effects (CATEs) with stable and accurate methods.
method R-Learning with inverse-variance weights (IVWs) for pseudo-outcome regression.
result IVWs improve the stability and accuracy of CATE estimation.
The study assesses external validity by evaluating worst-case treatment effects across subpopulations.
problem Underrepresentation of marginalized groups and limited study populations.
method Develops a semiparametrically efficient estimator for worst-case treatment effects (WTE) and uses cross-fitting to guard against brittle findings.
result The proposed framework guards against invalid findings due to unanticipated population shifts.
New approach for pricing evaluation improves on existing methods.
problem Improving off-policy evaluation for personalized pricing.
method Balanced policy evaluation framework with worst-case optimization.
result Empirical advantage over existing methods in pricing applications.
A framework for private causal effect estimation without structural assumptions.
problem Estimating causal effects from private observational data.
method Model-agnostic framework that privatizes predictions and aggregation steps.
result Maintains competitive performance under realistic privacy budgets.
Efficient policy learning from observational data using weighted classification reductions.
problem Efficient policy evaluation does not necessarily lead to efficient estimation of policy parameters.
method Proposed an estimation approach based on generalized method of moments, efficient for policy parameters.
result Demonstrated empirical efficiency and regret benefits of a proposed method.
Causal inference from observational data is hard due to discontinuous causal effects.
problem Causal inference from observational data is hard due to discontinuous causal effects.
method The problem is tackled by showing that many standard point estimates can be read as point summaries of multimodal distributions over the space of structural causal models.
result Many standard point estimates can be discontinuous summaries, while explicit posterior means and medians are continuous.
Differentially private method for estimating individualized treatment rules.
problem Estimating individualized treatment rules while preserving privacy.
method Differentially private two-stage empirical risk minimization (DP-2ERM).
result Improved privacy-utility trade-off demonstrated through simulations and applications.
Paper addresses regret minimization and inference in high-dimensional online decision-making.
problem Regret minimization and statistical inference in high-dimensional online decision-making.
method Integrates ε-greedy bandit algorithm with hard thresholding for sparse bandit parameters and debiasing method for inference.
result Achieves either O(T1/2) regret or O(T1/2)-consistent inference, with trade-off between exploration and exploitation. C-Learner improves stability of plug-in estimators for causal inference.
problem Limited overlap between treatment and control groups leads to unstable estimates.
method Constrained learning framework that achieves stability and asymptotic properties.
result Constrained learning produces stable estimates with desirable asymptotic properties.
The paper resolves the paradox of using unlabeled data for treatment effect estimation.
problem Using unlabeled data to estimate propensity scores for treatment effect estimation.
method Proposes a simple procedure to reconcile the use of estimated propensity scores with the advice to use true propensity scores.
result Direct regression may be preferable to inverse-propensity weighting in many circumstances.
This paper improves credit line impact analysis by considering spending as a distribution.
problem Previous studies on credit lines' impact on spending have overlooked the distributional nature of spending.
method Developed a distribution-valued estimator framework to extend existing real-valued estimators.
result Credit lines positively influence spending across all quantiles, but more towards luxuries as they increase.
The paper addresses statistical inference for online decision-making in a contextual bandit setting.
problem Understanding the performance of reward models in online decision-making with contextual information.
method The paper uses the contextual bandit framework with a linear reward model and the ε-greedy policy to address the exploration-exploitation dilemma. It employs the martingale central limit theorem and inverse propensity score weighting to establish asymptotic normality of parameter estimators. result The online ordinary least squares estimator and the online weighted least squares estimator are asymptotically normal, providing insights into the performance of the reward model.
Algorithm identifies best arm with biased proxy and selective ground truth audits.
problem Fixed-confidence best-arm identification with biased proxy and selective ground truth.
method Propensity-weighted estimator and adaptive auditing algorithm.
result Plug-in Neyman rule achieves near-oracle audit efficiency.
We present a new approach to the problems of evaluating and learning personalized decision policies from observational data of past contexts, decisions, and outcomes. Only the outcome of the enacted decision is available and the historical policy is unknown. These problems arise in personalized medicine using electroni…
Proposes a new IPW-based ranking metric for two-sided markets.
problem Addressing bias in implicit user feedback in two-sided markets.
method Extends IPW estimator to two-sided markets, addressing position bias.
result Proposed estimator is unbiased for ground-truth ranking metric.
DOPE efficiently estimates ATE with complex covariates.
problem Efficient estimation of ATE from complex covariates.
method Proposed DOPE framework for efficient adjustment.
result DOPE retains efficiency even with highly predictive covariates.
New method uses online learning to improve AIPW estimators for adaptively collected data.
problem Estimating treatment effects with adaptively collected data.
method Online learning to minimize sequentially weighted estimation error.
result Local minimax lower bound shows optimality of AIPW estimator.
New method reduces confidence interval sizes for causal inference.
problem Inaccurate propensity scores and extreme scores cause large confidence intervals.
method Data-dependent Coarse IPW (CIPW) estimators.
result Robust CIPW estimators reduce confidence interval sizes to ε+1/√n.
AM-PPI uses multiple predictors to reduce label cost in healthcare AI.
problem Reduces label cost in post-deployment monitoring of healthcare AI.
method Combines model predictions with a small labeled sample, routing each instance to a cost-appropriate subset of predictors.
result Produces narrower confidence intervals than single-predictor methods.
Algorithm reduces audit costs by identifying best service configurations from biased textual evidence.
problem Designing service systems from textual evidence requires accurate selection despite biased automated scoring.
method Developed PP-LUCB algorithm combining LLM scores and selective audits to minimize costs.
result Correctly identified the best model in 40/40 trials with 90% cost reduction.
Proposes DR algorithms for distributionally robust off-policy evaluation and learning.
problem Sensitive to environment distribution shifts in offline observational data.
method Doubly robust and distributionally robust approaches for OPE/L.
result Achieves semiparametric efficiency and fast regret rate.
Proposes a meta-learning method to improve recommender systems with biased feedback.
problem Learning from biased feedback in recommender systems.
method Asymmetric tri-training framework for meta-learning, using three predictors.
result Minimizes the upper bound of true performance metric, improving robustness to selection bias.
The paper addresses bias in fraud detection models by improving label recovery in payment networks.
problem Systematic bias in chargeback labels in payment networks.
method Formalizes the observation pipeline as a sequential missing-data problem with three stages and a corruption layer. Constructs the Sequential Triply Robust (STR) estimator to correct for all four impairments simultaneously.
result Achieves the semiparametric efficiency bound and provably dominates naive chargeback-based training in mean squared error.
New algorithm learns policies without uniform overlap assumption.
problem Learning optimal policies from non-uniformly collected data.
method Pessimistic Policy Learning (PPL) using lower confidence bounds.
result Efficient policy learning for adaptively collected data.
Study evaluates impact of probabilistic identity data in lookalike targeting campaigns.
problem Evaluate the impact of probabilistic identity data in lookalike targeting campaigns.
method Employ off-policy techniques to evaluate without risking large ad spend or A/B tests.
result Significant lift in conversion rate with identity-powered lookalikes.