Proposes methods for learning optimal dynamic treatment regimes robust to unconfoundedness violations.
problem Estimating optimal dynamic treatment regimes using historical observational data when unconfoundedness is violated.
method Utilizes proximal causal inference framework to propose three nonparametric identification methods, a (K+1)-robust method, and establish a semiparametric efficiency bound.
result Establishes the (K+1)-robust method for learning optimal dynamic treatment regimes, validating its efficiency and multiple robustness through numerical experiments.
Method uses semi-supervised learning to estimate optimal treatment regimes from medical records.
problem Estimating optimal treatment regimes from electronic medical records.
method Imputation-based semi-supervised method using unlabeled data.
result Proposed method yields more efficient estimators of optimal treatment regimes.
Optimal treatment regime uses proxy variables to improve decision-making.
problem Insufficient covariates in observational data lead to confounding issues.
method Proximal causal inference framework and outcome/treatment confounding bridges.
result The proposed optimal treatment regime outperforms existing ones.
The estimation of optimal treatment regimes is of considerable interest to precision medicine. In this work, we propose a causal k-nearest neighbor method to estimate the optimal treatment regime. The method roots in the framework of causal inference, and estimates the causal treatment effects within the nearest neig…
There is a fast-growing literature on estimating optimal treatment regimes based on randomized trials or observational studies under a key identifying condition of no unmeasured confounding. Because confounding by unmeasured factors cannot generally be ruled out with certainty in observational studies or randomized tri…
We propose a new procedure for inference on optimal treatment regimes in the model-free setting, which does not require to specify an outcome regression model. Existing model-free estimators for optimal treatment regimes are usually not suitable for the purpose of inference, because they either have nonstandard asympto…
Develops a method to estimate personalized treatment regimes from summary statistics.
problem Estimating optimal treatment regimes for a target population when individual-level data is unavailable.
method A weighting framework that tailors a treatment regime for the target population using summary statistics.
result Consistent and asymptotically normal estimator for optimal treatment regimes.
Dynamic treatment regimes are of growing interest across the clinical sciences as these regimes provide one way to operationalize and thus inform sequential personalized clinical decision making. A dynamic treatment regime is a sequence of decision rules, with a decision rule per stage of clinical intervention; each de…
New method evaluates personalized treatment in critical care, robust to death.
problem Truncation by death in critical care makes traditional DTR evaluation ineffective.
method Principal stratification-based approach, focusing on always-survivor value function, with a semiparametrically efficient, multiply robust estimator.
result Demonstrates robustness and efficiency of the method for personalized treatment optimization.
Optimal treatment regimes (OTR) are individualised treatment assignment strategies that identify a medical treatment as optimal given all background information available on the individual. We discuss Bayes optimal treatment regimes estimated using a loss function defined on the bivariate distribution of dichotomous po…
Method learns optimal treatment sequences from observational data.
problem Optimal dynamic treatment regimes for public policies and medical interventions.
method Doubly robust classification-based approach via backward induction.
result Achieves optimal convergence rate of n^(-1/2) for welfare regret.
The vision for precision medicine is to use individual patient characteristics to inform a personalized treatment plan that leads to the best healthcare possible for each patient. Mobile technologies have an important role to play in this vision as they offer a means to monitor a patient's health status in real-time an…
New method for finding optimal treatment regimes in medical settings with time-varying unobserved factors.
problem Finding optimal treatment regimes in medical settings with time-varying unobserved factors.
method Extend Dynamic Treatment Regimes (DTRs) to Ambiguous Dynamic Treatment Regimes (ADTRs), connect to Ambiguous Partially Observable Mark Decision Processes (APOMDPs), and develop Reinforcement Learning methods.
result Established theoretical results for learning methods, including consistency and asymptotic normality.
Develops methods to learn optimal treatment regimes using causal tree methods.
problem Lack of methods for estimating treatment effects and handling complex patient data.
method Causal tree and causal forest methods for estimating heterogeneous treatment effects.
result Outperforms state-of-the-art baselines in cumulative regret and percentage of optimal decisions.
Proposes a new method for dynamic treatment regimes that improves sample efficiency and stability.
problem Challenges in estimating optimal treatments for individuals with dynamic decision-making stages.
method Focuses on prioritizing alignment between observed and optimal treatment trajectories across decision stages.
result Improves sample efficiency and stability of IPWE-based methods by relaxing the alignment requirement.
The paper develops methods to estimate optimal treatment sequences under policy constraints.
problem Estimating the best sequence of treatments over multiple stages for individuals.
method Empirical welfare maximization approach, solving treatment assignment sequentially or simultaneously.
result Established convergence rates and upper bounds for estimation methods.
Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.
problem Challenges in learning optimal dynamic treatment regimes with large intervention options and infinite time horizon.
method Proximal Temporal consistency Learning (pT-Learning) framework for adaptively adjusting between deterministic and stochastic policies.
result Minimax estimator avoids double sampling issue and can incorporate off-policy data.
The application of existing methods for constructing optimal dynamic treatment regimes is limited to cases where investigators are interested in optimizing a utility function over a fixed period of time (finite horizon). In this manuscript, we develop an inferential procedure based on temporal difference residuals for …
Develops a novel approach for estimating optimal DTRs with multicategory treatments and censored data.
problem Estimating optimal treatment regimes for chronic diseases with censored data.
method Angle-based multicategory classification algorithm for maximizing conditional survival function.
result The proposed method outperforms existing approaches in maximizing conditional survival function.
Q-learning adapted to find near-equivalent treatment strategies.
problem Finding optimal treatment sequences in dynamic treatment regimes.
method Introducing a worst-value tolerance criterion to find sets of near-equivalent policies.
result Constructs families of near-equivalent treatment strategies.
SAFER improves personalized treatment recommendations for dynamic clinical contexts.
problem Personalized treatment optimization in evolving clinical contexts with safety concerns.
method Integrates structured EHR and clinical notes, uses conformal prediction for safe recommendations.
result SAFER outperforms state-of-the-art baselines in recommendation metrics and mortality rates.
Variable selection for optimal treatment regime in a clinical trial or an observational study is getting more attention. Most existing variable selection techniques focused on selecting variables that are important for prediction, therefore some variables that are poor in prediction but are critical for decision-making…
Comment on entropy learning for dynamic treatment regimes.
problem Evaluating dynamic treatment regimes using entropy loss.
method Optimization-based alternative to IPW estimate.
result Suggests optimization-based approach for evaluation.
This paper presents the first deep reinforcement learning (DRL) framework to estimate the optimal Dynamic Treatment Regimes from observational medical data. This framework is more flexible and adaptive for high dimensional action and state spaces than existing reinforcement learning methods to model real-life complexit…
Proposes a new Bayesian learning method for optimal treatment regimes.
problem Sub-optimal policies in offline data due to lack of exploration.
method Integrates pessimism principle with Thompson sampling and Bayesian machine learning.
result Derives a credible set that uniformly lower bounds the optimal Q-function.
POLAR optimizes treatment strategies in dynamic settings with statistical guarantees.
problem Optimizing sequential decisions in dynamic treatment regimes with robustness and statistical guarantees.
method Pessimistic model-based approach estimating transition dynamics and incorporating uncertainty penalties.
result Offers statistical and computational guarantees, including finite-sample bounds on policy suboptimality.
A new algorithm learns optimal personalized treatment plans online with low regret.
problem Learning optimal dynamic treatment regimes in an online setting.
method Developed a novel algorithm balancing exploration and exploitation for rate-optimal regret.
result Guaranteed rate-optimal regret for linear transition and reward models.
Develops methods for personalized treatment decisions in the presence of unmeasured factors.
problem Personalized treatment decisions in the presence of unmeasured confounding.
method Proximal learning approaches to estimate optimal individualized treatment regimes (ITRs).
result Established identification results for different classes of ITRs, improving decision-making value function.
Proposes a method to estimate personalized treatments from high-dimensional data.
problem Estimating individualized treatment regimes (ITRs) from high-dimensional covariates.
method Directly targets the contrast between potential outcomes, using dimension-reduced outcome-weighted learning.
result Achieves universal consistency, converging to the Bayes risk under mild conditions.
Develops framework for estimating and improving DTRs with time-varying IV in the presence of unmeasured confounding.
problem Estimating DTRs from observational data with unmeasured confounding.
method Time-varying instrumental variable (IV) framework for estimating and improving DTRs.
result IV-optimal and IV-improved DTRs perform better than DTRs assuming no unmeasured confounding.
A treatment regime is a function that maps individual patient information to a recommended treatment, hence explicitly incorporating the heterogeneity in need for treatment across individuals. Patient responses are dichotomous and can be predicted through an unknown relationship that depends on the patient information …
New framework estimates treatment effects in extreme data.
problem Hindered by unavailability of counterfactual outcomes and rarity of extreme data.
method Proposes a new framework based on extreme value theory.
result Quantifies treatment effects using tail decay rates of potential outcomes.
The field of precision medicine aims to tailor treatment based on patient-specific factors in a reproducible way. To this end, estimating an optimal individualized treatment regime (ITR) that recommends treatment decisions based on patient characteristics to maximize the mean of a pre-specified outcome is of particular…
RL algorithms with medical integration improve personalized treatment recommendations.
problem Developing effective personalized treatment strategies for chronic diseases.
method Integrating medical knowledge into RL algorithms for DTR.
result Enhanced treatment recommendations with increased confidence.
New method estimates optimal personalized treatment rules from mixed data sources.
problem Combining RCT and observational data for personalized treatment rules.
method Doubly robust estimator for value function, maximizing within pre-specified ITR class.
result Consistent and asymptotically normal optimal value estimator with N−1/3 rate of convergence. New method for estimating treatment effects without complex propensity models.
problem Estimating treatment effects in dynamic treatment regimes.
method Recursive Riesz representer estimation for de-biasing corrections.
result Directly estimates de-biasing corrections without auxiliary models.
We develop and evaluate tolerance interval methods for dynamic treatment regimes (DTRs) that can provide more detailed prognostic information to patients who will follow an estimated optimal regime. Although the problem of constructing confidence intervals for DTRs has been extensively studied, prediction and tolerance…
The paper introduces a new method for estimating optimal policies in dynamic treatment regimes using information geometry.
problem Estimating optimal policies in dynamic treatment regimes.
method Minimum information divergence method based on γ-power divergence. result The γ-power divergence method effectively seeks the optimal policy by vanishing the divergence between policy-equivalent Q-functions. Sensitivity analysis for individualized effects in OTRs with binary risk factors.
problem Addressing omitted confounding in individualized effects of OTRs.
method Simulation-based sensitivity analysis to simulate unmeasured confounders.
result Benchmarking the strength of omitted confounding for binary risk factors.
Develops a new RL algorithm for medical treatment regimes.
problem Optimal dose determination in continuous action environments.
method Quasi-optimal learning algorithm for near-optimal actions.
result Guaranteed convergence and effectiveness in real applications.
Develops methods to identify and estimate causal effects with instrumental variables.
problem Causal inference with confounded treatment assignment and unobserved variables.
method General nonparametric causal framework, debiased machine learning, semiparametric theory.
result Consistent and asymptotically normal estimators for average treatment effect.
Paper proves optimality of doubly robust estimators for treatment effects.
problem Estimating treatment effects in causal inference.
method Structure-agnostic framework of statistical lower bounds, using non-parametric regression and classification oracles.
result Doubly robust estimators are statistically optimal for ATE and ATT.
Study best arm identification with contextual info, achieving optimal misidentification probability.
problem Identify the best treatment arm with minimal misidentification probability in a small gap scenario.
method Developed RS-AIPW strategy that matches lower bound of misidentification probability in the small-gap regime.
result RS-AIPW strategy is asymptotically optimal for best arm identification.
In order to identify important variables that are involved in making optimal treatment decision, Lu et al. (2013) proposed a penalized least squared regression framework for a fixed number of predictors, which is robust against the misspecification of the conditional mean model. Two problems arise: (i) in a world of ex…
Deep Bayesian models estimate causal effects for dynamic treatment regimes over long follow-up times.
problem Challenges in causal effect estimation for dynamic treatment regimes with long follow-up times.
method Combining outcome regression models with deep Bayesian models for high-dimensional features.
result Stable and accurate dynamic causal effect estimation from observational data, especially with long-term follow-up.
Study optimal adjustment sets for causal policies with hidden variables.
problem Estimating dynamic treatment regimes with hidden variables.
method Developed criteria for graphs without hidden variables to compare estimators, extended to dynamic policies and hidden variables.
result Existence and computation of optimal minimal and globally optimal adjustment sets.
LUQ-Learning adapts Q-learning for healthcare decisions considering patient preferences.
problem Optimizing treatment decisions for multivariate outcomes based on individual preferences.
method Latent Utility Q-Learning (LUQ-Learning) framework that adapts Q-learning for composite outcomes.
result LUQ-Learning achieves highly competitive performance compared to alternative methods in simulations.
We study the problem of learning to choose from m discrete treatment options (e.g., news item or medical drug) the one with best causal effect for a particular instance (e.g., user or patient) where the training data consists of passive observations of covariates, treatment, and the outcome of the treatment. The standa…