Machine learning boosts RCT efficiency by controlling type I error and improving statistical power.
problem Improving statistical efficiency in RCTs with complex covariate adjustments.
method Machine learning-assisted adjustment under Rosenbaum's framework for exact tests.
result The proposed method robustly controls type I error and significantly boosts statistical efficiency.
The paper compares methods for estimating heterogeneous treatment effects using multiple randomized trials.
problem Estimating heterogeneous treatment effects reliably and precisely with a single dataset is challenging.
method Non-parametric approaches for estimating heterogeneous treatment effects using data from multiple trials.
result Methods that directly allow for heterogeneity of the treatment effect across trials perform better than those that do not.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
problem Minimizing subgroup bias in randomized controlled trials.
method Wasserstein Homogeneity Partition (WHOMP) method.
result WHOMP optimally minimizes type I and type II errors in trials.
New estimator improves policy evaluation in resource allocation RCTs.
problem Difficulty in evaluating policies optimizing limited resource allocation through RCTs.
method Proposes a novel estimator involving retrospective reshuffling of participants across experimental arms.
result The new estimator provides more accurate policy evaluations than common methods.
New methods improve subgroup analysis in trials with limited data.
problem Limited sample sizes in subgroup analyses of randomized controlled trials.
method Two TMLEs that borrow information from non-subgroup participants.
result Improved precision in subgroup-specific treatment effect estimates.
Adaptive Prespecification improves precision in randomized trials.
problem Selecting optimal covariates for precision in randomized trials.
method Adaptive Prespecification using V-fold cross-validation and influence curve-squared loss function.
result Substantial gains in precision, equivalent to 20-43% reductions in sample size for the same power.
The paper evaluates index-based allocation policies using data from randomized control trials.
problem Evaluating index-based allocation policies in resource-scarce scenarios.
method Using data from randomized control trials, the paper introduces an efficient estimator and methods for computing asymptotically correct confidence intervals.
result Valid statistical conclusions can be drawn for index-based allocation policies.
New method detects biomarker-treatment interactions in clinical trials.
problem Detecting interactions between high-dimensional biomarkers and treatments in randomized trials.
method Two-stage penalized regression screening using ridge regression for multivariate screening.
result Ridge regression screening provides greater power than traditional methods in correlated data.
Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform randomization. We show that this procedure can be highly sub-optimal (in terms of learnin…
This study evaluates subgroup analysis methods for time-to-event outcomes in randomized controlled trials.
problem Identifying subgroups of good responders in non-significant randomized controlled trials.
method Evaluation of several subgroup analysis algorithms for time-to-event outcomes using synthetic and semi-synthetic data.
result Provides a new synthetic and semi-synthetic data generation process and an open-source Python package for benchmarking.
A-TMLE estimates ATE from RCT and RWD, achieving super-efficiency.
problem Estimating ATE from RCT and RWD data.
method Adaptive-TMLE framework for decomposing and estimating ATE.
result A-TMLE is root-n consistent and asymptotically normal, achieving super-efficiency.
Framework for estimating treatment effects using external control data.
problem Improving efficiency in estimating average treatment effects (ATE) in hybrid trials.
method Developed a formal causal inference framework based on exchangeability assumptions and graphical criteria. Proposed estimators and efficient doubly-robust methods.
result Established finite-sample performance and demonstrated application to spinal muscular atrophy trial.
Machine learning improves learning and memory retention by optimizing study sessions.
problem Improving learning and memory retention methods for factual material.
method Large-scale randomized controlled trial with machine learning optimization of study sessions.
result Study sessions optimized with machine learning lead to 67% longer retention and 50% higher return rate.
Participants enrolled into randomized controlled trials (RCTs) often do not reflect real-world populations. Previous research in how best to translate RCT results to target populations has focused on weighting RCT data to look like the target data. Simulation work, however, has suggested that an outcome model approach …
Causal ML methods failed to validate their personalized treatment effects in two large trials.
problem Validating causal machine learning methods for personalized treatment effects in precision medicine.
method Assessed 17 mainstream causal heterogeneity ML methods using two large randomized controlled trials.
result None of the ML methods reliably validated their performance, internal or external, showing significant discrepancies between training and test data.
Digital twins improve single-arm trials by providing robust treatment effect estimates.
problem Lack of control arms in single-arm trials limits their gold-standard evidence.
method Outcome-model-based synthetic controls using machine learning models trained on historical data.
result Digital twins offer more robust treatment effect estimates and principled corrections.
Two-stage TMLE reduces bias and improves efficiency in CRTs.
problem Differential outcome measurement and imbalance in baseline predictors in CRTs.
method Two-stage targeted minimum loss-based estimator (TMLE) to adjust for baseline covariates.
result Our approach nearly eliminates bias due to differential outcome measurement.
G-computation improves clinical trial power with machine learning.
problem Balancing prognostic factors in randomized trials to prevent near-confounders.
method G-computation with penalized models (Lasso, Elasticnet) and algorithm-based methods (neural network, SVM, super learner).
result G-computation with Elasticnet and splines reduces variance and increases power in RCTs.
New TTP framework fuses control arms while controlling Type-I error.
problem Bias in borrowing control data from previous trials.
method Kernel two-sample testing via MMD and equivalence testing.
result Higher power than standard TTP methods while maintaining error control.
New framework for adaptive clinical trials to address real-world challenges.
problem Real-world challenges in post-regulatory clinical trials.
method RFAN framework integrating regulatory constraints and treatment policy value.
result Empirical evaluation of RFAN's performance.
The P300 event-related potential (ERP), evoked in scalp-recorded electroencephalography (EEG) by external stimuli, has proven to be a reliable response for controlling a BCI. The P300 component of an event related potential is thus widely used in brain-computer interfaces to translate the subjects' intent by mere thoug…
DARTS optimizes covariate selection in trials with limited data.
problem Limited budget for high-dimensional pretreatment data.
method Dynamic Adaptive Rerandomization via Thompson Sampling (DARTS).
result DARTS efficiently concentrates budget on informative features.
New approach estimates treatment effects from decentralized data.
problem Estimating treatment effects from multiple studies with limited data.
method Three classes of ATE estimators derived from Plug-in G-Formula.
result Asymptotic variance of estimators for linear models derived.
New methods improve causal inference generalization using trial and observational data.
problem Limited trial data makes generalizing causal inferences to target populations statistically infeasible.
method Develops algorithms that combine trial and observational data to estimate complex nuisance functions.
result Improves generalization of causal inferences when the additional observational study is high-quality.
MEC-Cox: A Machine-Learning-Assisted Generalized Entropy Calibration Method for Estimating ATT Marginal Hazard-Ratio
problem Estimating ATT marginal hazard-ratio in externally controlled survival trials
method Machine-learning-assisted generalized entropy calibration for IPW Cox regression
result Reduces bias, increases efficiency, and improves coverage
New method uses latent variables to estimate treatment effects from single-arm trials.
problem Estimating treatment effects from single-arm trials due to lack of external control groups.
method Latent-variable modeling with amortized variational inference for patient matching and direct effect estimation.
result Improved performance in direct treatment effect estimation and effect estimation via patient matching compared to previous methods.
Improves trial efficiency by adjusting for historical prognostic scores.
problem Reducing statistical uncertainty in randomized trial estimates.
method Linear covariate adjustment using a prognostic model trained on historical data.
result Prognostic covariate adjustment achieves minimum variance and reduces mean-squared error.
New method uses observational data to improve trial design efficiency.
problem Scarce randomized controlled trials; inefficiency of using observational data.
method Active Residual Learning, R-Design framework, R-EPIG criterion.
result Efficiently estimating residuals to correct observational bias improves trial design.
Syntax designs adaptive trials for subpopulations with potential benefits.
problem Identifying subpopulations with positive treatment effects in diverse patient populations.
method Adaptive patient recruitment and synthetic control estimation.
result Syntax outperforms conventional trial designs in identifying beneficial subpopulations.
Generative AI models improve clinical trial data by generating survival outcomes.
problem Generating valid survival outcomes for clinical trials with synthetic data.
method A variational autoencoder (VAE) that jointly generates mixed-type covariates and survival outcomes.
result The method outperforms GAN baselines on fidelity, utility, and privacy metrics.
New methods optimize personalized treatment assignment in trials with many arms.
problem Poor performance of standard methods in trials with many treatment arms.
method Regularized and clustered joint assignment forest algorithm.
result Gains in predicting arm-wise outcomes and utility gains from personalization.
Observed events in recommendation are consequence of the decisions made by a policy, thus they are usually selectively labeled, namely the data are Missing Not At Random (MNAR), which often causes large bias to the estimate of true outcomes risk. A general approach to correct MNAR bias is performing small Randomized Co…
Study on the probability of immunity and its bounds.
problem Estimating the probability of immunity and its bounds.
method Derive necessary and sufficient conditions for non-immunity and ε-bounded immunity; introduce indirect immunity; propose sensitivity analysis.
result Estimate the probability of benefit and produce tighter bounds of the probability of benefit.
Paper introduces active and passive causal inference techniques.
problem Causal inference in machine learning.
method Categorizes causal inference techniques into active and passive approaches.
result Describes and discusses various causal inference methods.
Machine learning reduces variance in online experiment results.
problem Reducing variance in randomized controlled trials.
method Machine learning regression-adjusted treatment effect estimator (MLRATE).
result MLRATE reduces estimator variance by over 70% in A/A tests.
The paper proposes a method to find subgroups with significant treatment effects in noisy data.
problem Estimating the causal effects of interventions on noisy outcomes.
method A machine-learning method specifically optimized for finding subgroups with significant effects, designed to maximize the probability of obtaining a statistically significant positive treatment effect.
result The proposed method yields higher power in detecting subgroups affected by the treatment compared to standard tree-based tools.
Studies across many disciplines have shown that lexical choice can affect audience perception. For example, how users describe themselves in a social media profile can affect their perceived socio-economic status. However, we lack general methods for estimating the causal effect of lexical choice on the perception of a…
New study shows non-adaptive trials can be outperformed by adaptive designs in treatment selection.
problem Determining the best allocation of resources in clinical trials.
method Analysis of batched arm elimination designs and comparison with completely randomized trials.
result Simple adaptive designs universally and strictly dominate non-adaptive completely randomized trials for at least three treatment arms.
QR-learner estimates individual treatment effects using external data.
problem Limited power to detect individual treatment effects in randomized trials.
method Model-agnostic learner that estimates conditional average treatment effects (CATE) using external data.
result QR-learner reduces mean squared error and can recover true CATE.
Framework tests CATE homogeneity across trials and evaluates confounding.
problem Assessing treatment effect consistency across randomized and observational studies.
method Leverages multiple randomized trials to test CATE homogeneity and compares with observational data.
result Identifies potential confounding and effect heterogeneity in treatment effects.
This study quantifies uncertainty in comparing treatments using RCTs with before-and-after measures.
problem Uncertainty in comparing treatments using RCTs with before-and-after measures.
method New statistical modeling principle called ETZ enables counterfactual uncertainty quantification (CUQ) in RCTs with Before-and-After Repeated Measures.
result CUQ typically has lower variability than factual uncertainty quantification and can be achieved in RCTs.
Combines trial and observational data to improve policy evaluation.
problem External validity of randomized trial results in target populations.
method Uses covariate data to model trial sampling and certifies policy evaluations.
result Valid trial-based policy evaluations under model miscalibration.
A new method improves treatment effect inferences in RCTs by adjusting for covariates and heteroskedasticity.
problem Improving treatment effect inferences in RCTs with efficient and powerful methods.
method Weighted Prognostic Covariate Adjustment Method (Weighted PROCOVA) for heteroskedasticity.
result The method reduces variance, maintains Type I error rate, and increases test power for treatment effect.
A new method uses randomized trials to estimate the strength of unobserved confounding.
problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.
Large dataset released for ITE and UM research.
problem Estimating causal impact of actions in various sectors.
method Release of a large dataset, formalization of UM, synthetic response surfaces, heterogeneous treatment assignment.
result Validation of ITE prediction and UM methods with high statistical significance.
Proposes a method to correct for covariate shift in meta-analysis of randomized trials.
problem Invalidation of standard IPD meta-analysis due to covariate shift across studies.
method Placebo-anchored transport framework that treats source-trial outcomes as proxy signals and target-trial placebo outcomes as gold labels.
result Yields target-identified effect estimates in connected targets and a principled screen--then--transport procedure in disconnected targets.
Customer scoring models are the core of scalable direct marketing. Uplift models provide an estimate of the incremental benefit from a treatment that is used for operational decision-making. Training and monitoring of uplift models require experimental data. However, the collection of data under randomized treatment as…
Method controls treatment risk in learning beneficial allocations.
problem Learning beneficial treatment allocations with risk control in precision medicine.
method Proposes a certifiable learning method that controls treatment risk with finite samples in the partially identified setting.
result Illustrates method using both simulated and real data.