New method for estimating parameters in inverse problems using double robustness.
problem Estimating parameters defined as linear functionals of solutions to linear inverse problems.
method Source condition double robust inference method that uses iterated Tikhonov regularized adversarial estimators.
result Asymptotic normality of the parameter of interest as long as either the primal or dual inverse problem is sufficiently well-posed.
New method improves robustness of double robust estimators under complete misspecification.
problem Improper performance of double robust estimators when all nuisance functions are misspecified.
method DR+ACC, an adaptive correction clipping method.
result DR+ACC ensures bounded error and maintains semiparametric efficiency.
Proposes fair and robust methods for estimating treatment effects.
problem Estimating treatment effects while maintaining fairness.
method Simple, nonparametric framework with fairness constraints.
result Estimators are double robust and characterize welfare trade-offs.
Tests validity of DML estimators without assumptions.
problem Validating DML estimators without making assumptions.
method Develops tests to falsify assumptions for DML estimators.
result Falsifies assumptions for DML estimators with non-trivial power.
Paper introduces GDR-learners for estimating potential outcomes from observational data.
problem Lack of theoretical property of general Neyman-orthogonality in deep generative models.
method Develops flexible GDR-learners based on various deep generative models.
result GDR-learners possess quasi-oracle efficiency and rate double robustness, asymptotically optimal.
Bayesian method for estimating ATE with robustness to model misspecification.
problem Estimating average treatment effects under unconfoundedness.
method Double robust Bayesian inference using adjusted prior and posterior distributions.
result Bayesian credible sets form asymptotically exact confidence intervals.
Direct learning framework for integrating multi-source causal data.
problem Conditional average treatment effects inference from heterogeneous data.
method Direct learning framework, double robustness, causal information-aware weighting function.
result Effective causal data fusion in both homogeneous and heterogeneous scenarios.
Consider the case that one observes a single time-series, where at each time t one observes a data record O(t) involving treatment nodes A(t), possible covariates L(t) and an outcome node Y(t). The data record at time t carries information for an (potentially causal) effect of the treatment A(t) on the outcome Y(t), in…
The paper tackles fairness in data and algorithms, expanding on prior work.
problem Discrimination and disparate treatment in data and algorithms.
method Targeted learning for nonparametric inference of fairness in the data generating process.
result Derivation and validation of estimators for fairness metrics like demographic parity and equal opportunity.
When training a machine learning model with observational data, it is often encountered that some values are systemically missing. Learning from the incomplete data in which the missingness depends on some covariates may lead to biased estimation of parameters and even harm the fairness of decision outcome. This paper …
Paper tackles efficient policy gradient estimation from off-policy data.
problem Estimating policy gradients from off-policy data is challenging and inefficient.
method Derives asymptotic lower bounds, proposes a meta-algorithm with 3-way robustness, and establishes convergence guarantees.
result Meta-algorithm achieves the lower bound on mean-squared error without parametric assumptions.
New estimators for causal effects in DAGs with hidden variables, addressing computational and statistical challenges.
problem Estimating causal effects in DAGs with hidden variables beyond traditional criteria.
method Introduces novel one-step corrected plug-in and targeted minimum loss-based estimators for causal effects in DAGs with hidden variables.
result Root-n consistent causal effect estimates with desirable statistical properties.
We develop methods to approximate derivatives for causal inference problems using data.
problem Estimating causal effects from data when distributions are not known.
method Constructive algorithm approximating Gateaux derivatives via finite differencing.
result Derives conditions for finite-difference approximations to preserve statistical benefits.
New method for estimating mean in SS inference with selection bias and decaying overlap.
problem Estimating mean in SS inference with selection bias and decaying overlap.
method Double Robust Semi-Supervised (DRSS) mean estimator.
result Consistent estimation of mean with correct specification of outcome or propensity score model.
A theorem for debiasing machine learning with finite sample guarantees.
problem Calculating confidence intervals for machine learning functionals.
method Debiased machine learning based on bias correction and sample splitting.
result Nonasymptotic debiased machine learning theorem with finite sample guarantees.
Proposes CCME framework for estimating heterogeneous treatment effects.
problem Estimating heterogeneous treatment effects in complex distributions.
method Embeds conditional distributions into RKHS, develops meta-estimators for CCME.
result Establishes finite-sample convergence rates and double robustness for CCME estimators.
Proposes a scalable method for counterfactual prediction using machine learning.
problem De-bias causal estimators with high-dimensional data in observational studies.
method Uses entropy balancing to learn weights minimizing Jensen-Shannon divergence, leading to robust counterfactual predictions.
result Consistent causal estimation if either propensity score or outcome model is correctly specified.
The paper addresses bias in survival analysis due to informative censoring.
problem Bias in treatment effect estimates due to informative censoring in survival analysis.
method Assumption-lean framework using partial identification to derive bounds on CATE.
result Proposes a meta-learner, SurvB-learner, to estimate bounds on CATE.
Paper develops a new estimator for dynamic treatment effects in high-dimensional settings.
problem Time-varying confounding and model misspecification in estimating dynamic treatment effects.
method Sequential model doubly robust estimator with moment-targeting estimates.
result Root-N inference achieved under model misspecification, even with high-dimensional covariates.
The paper investigates how calibrating propensity scores improves DML estimates of average treatment effects.
problem Improving the accuracy of DML estimates in finite samples.
method Propensity score calibration within the Double/debiased machine learning framework.
result Calibrating propensity scores reduces the root mean squared error of DML estimates of average treatment effects in finite samples.
A new method improves estimation of COVID-19 vaccine effectiveness.
problem Estimating vaccine effectiveness under the test-negative design.
method A doubly robust estimator (TNDDR) using cross-fitting and machine learning.
result The TNDDR estimator is n \sqrt{n} n -consistent, asymptotically normal, and doubly robust. New methods improve off-policy evaluation for survival outcomes with censoring.
problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.
Proposes nAIPW for robust ATE estimation using neural networks.
problem Estimation of ATE with potential confounders and nonlinear relationships.
method Normalized AIPW (nAIPW) with neural networks and regularization.
result nAIPW maintains double-robustness and orthogonality properties.
New methods for estimating causal effects in hidden variable DAGs.
problem Estimating causal effects in models with hidden variables.
method Influence function based estimators for causal effects in hidden variable DAGs.
result Achieves semiparametric efficiency bounds for identifiable effects.
Directly estimates CQC, improving interpretability and accuracy.
problem Inability to model and interpret CQC due to inversion issue.
method Direct doubly robust estimation of CQC without inversion.
result Improved estimation accuracy and interpretability.
Automatic debiasing for causal and policy effects using Neural Nets and Random Forests.
problem Estimating causal and policy effects from high-dimensional or non-parametric regression functions.
method Automatic learning of Riesz representation using Neural Nets and Random Forests.
result Automatic debiasing method performs well compared to state-of-the-art algorithms.
We propose a semiparametric test to evaluate (i) whether different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. The test is a flexible robustness check f…
A new method improves efficiency in finding optimal personalized treatment rules.
problem Heteroscedasticity and misspecified treatment-free effect models affect optimal ITR estimation.
method E-Learning framework that accounts for covariate-treatment dependent variance of residuals.
result E-Learning framework improves efficiency of optimal ITR estimation.
TMLE improves unbiased estimation in public health studies.
problem Improving unbiased estimation in observational studies.
method Targeted Maximum Likelihood Estimation (TMLE) integrates machine learning and statistical theory.
result TMLE has been adopted by researchers worldwide, especially outside the US.
Adapts randomization for single unit time-series data for optimal treatment.
problem Statistical methods for precision medicine in single unit time-series data.
method Adaptive sequential design, nonparametric model, double robust structure, efficient influence function.
result Valid inference for mean target parameter based on single sample.
Observational cohort studies with oversampled exposed subjects are typically implemented to understand the causal effect of a rare exposure. Because the distribution of exposed subjects in the sample differs from the source population, estimation of a propensity score function (i.e., probability of exposure given basel…
A fundamental challenge in semi-supervised learning lies in the observed data's disproportional size when compared with the size of the data collected with missing outcomes. An implicit understanding is that the dataset with missing outcomes, being significantly larger, ought to improve estimation and inference. Howeve…
Proposes isotonic regression for calibrating Deep Cox models' survival probabilities.
problem Poor calibration of Deep Cox models' survival probabilities.
method Isotonic regression for post hoc calibration of Deep Cox models.
result Establishes favorable theoretical guarantees and demonstrates empirical effectiveness.
Unified framework for robust causal directionality in quantum systems under MNAR observation.
problem Determining causal directionality in quantum systems under MNAR observation.
method Integrates CVAE-based latent constraints, MNAR-aware selection models, GEE-stabilized regression, penalized empirical likelihood, and Bayesian optimization.
result Achieves lower bias and variance, near-nominal coverage, and superior quantum-specific diagnostics.
The paper addresses the gap between theoretical and practical confidence set widths in universal inference.
problem Inference procedures can be overly conservative, leading to wider confidence sets than expected.
method The authors identify the source of asymptotic conservativeness and propose a remedy based on studentization and bias correction.
result The proposed method achieves exact asymptotic coverage at the nominal 1 − α 1-α 1 − α level, even under model misspecification. Estimates causal effect of managed care plans on NYC Medicaid spending.
problem Generalizing causal estimates to a target population not well-represented by randomized studies.
method Conditional cross-design synthesis estimators combining randomized and observational data.
result Estimates causal effect of managed care plans on health care spending.
Paper tackles efficient evaluation of natural stochastic policies in offline RL.
problem Efficiency issues in evaluating natural stochastic policies due to unknown evaluation policy.
method Derive efficiency bounds for tilting and modified treatment policies, propose nonparametric estimators.
result Proposed estimators attain efficiency bounds under lax conditions and enjoy partial double robustness.
We study minimax methods for off-policy evaluation (OPE) using value functions and marginalized importance weights. Despite that they hold promises of overcoming the exponential variance in traditional importance sampling, several key problems remain: (1) They require function approximation and are generally biased. Fo…
Paper tackles moment estimation under covariate shift with a two-stage algorithm.
problem Estimating moments under covariate shift when source and target distributions differ.
method Proposes a two-stage algorithm: first, an optimal estimator for the source distribution; second, likelihood ratio reweighting for calibration.
result Achieves minimax optimal bound for moment estimation.
Paper proposes methods to reduce bias and variance in recommender systems.
problem Bias in recommender systems due to users' preferences.
method Proposes a principled approach to reduce bias and variance in DR methods, and a novel semi-parametric collaborative learning approach.
result The proposed methods outperform existing debiasing methods in both theory and experiments.
StableDR stabilizes doubly robust learning for biased recommendation data.
problem Data missing not at random in recommender systems.
method StableDR, a stabilized doubly robust learning approach.
result StableDR achieves bounded bias, variance, and generalization error.
Estimating linear, mean-square continuous functionals is a pivotal challenge in statistics. In high-dimensional contexts, this estimation is often performed under the assumption of exact model sparsity, meaning that only a small number of parameters are precisely non-zero. This excludes models where linear formulations…
C-Learner improves stability of plug-in estimators for causal inference.
problem Limited overlap between treatment and control groups leads to unstable estimates.
method Constrained learning framework that achieves stability and asymptotic properties.
result Constrained learning produces stable estimates with desirable asymptotic properties.
Estimating causal effects for survival outcomes in the high-dimensional setting is an extremely important topic for many biomedical applications as well as areas of social sciences. We propose a new orthogonal score method for treatment effect estimation and inference that results in asymptotically valid confidence int…
While model selection is a well-studied topic in parametric and nonparametric regression or density estimation, selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we propose a selective machine learning framework for making inferences about a fini…
Bayesian method corrects bias in treatment effect estimation.
problem Estimating treatment effects from observational data with high-dimensional nuisance parameters.
method Bayesian debiasing, targeted modeling, sample splitting.
result Marginal posterior for ATE satisfies Bernstein-von Mises theorem under correct nuisance model specification.
SCIENCE improves prediction intervals for individual causal effects.
problem Wide prediction intervals limit practical utility of causal inference.
method Surrogate-assisted conformal inference for efficient individual causal effects.
result SCIENCE produces more efficient prediction intervals for individual causal effects.
New method improves CATE model selection with optimal regret rates.
problem Nontrivial task of selecting accurate CATE models.
method Causal Q-aggregation using doubly robust loss.
result Achieves optimal oracle model selection regret rates of log(M)/n.