This work explains RL policies using causal models, revealing important patterns and failures.
problem Understanding why RL policies succeed or fail in complex, high-dimensional systems.
method Developed a nonlinear Causal Model Reduction framework to learn simplified causal models from RL policy actions and rewards.
result The approach can uncover important behavioral patterns and failure modes in trained RL policies.
End-to-end policy learning method improves CATE estimation.
problem Learning optimal treatment policies from partially observed data.
method Modified causal forest for policy learning.
result Maximizing policy value is equivalent to minimizing CATE.
AACE learns treatment policies from EHRs using annotations to improve accuracy.
problem Learning treatment policies from multimodal EHRs with bias and inefficiency.
method Annotation-assisted coarsened effects (AACE) method.
result AACE outperforms existing methods in predicting treatment benefit from multimodal EHRs.
Researchers use DT to transfer policies from one environment to another using causal reasoning.
problem Adapting to changes in environmental dynamics in reinforcement learning.
method Applying causal counterfactual reasoning to Decision Transformer (DT) architecture for policy transfer.
result DT successfully transfers a learned policy to new environments while retaining most of the reward.
The paper tackles policy learning in dynamic environments using causal methods.
problem Existing reinforcement learning algorithms assume static mechanisms, but real-world systems often have changing mechanisms.
method The paper introduces multi-environment contextual bandits and policy invariance to handle environmental shifts.
result An optimal invariant policy is guaranteed to generalize across environments under suitable assumptions.
ICIL learns policies invariant to multiple environments, improving generalization.
problem Learning policies from multiple environments leads to spurious correlations.
method ICIL learns invariant feature representations and a matching imitation policy.
result ICIL policies generalize better to unseen environments.
New methods use vector search and nearest-neighbor matching for policy learning in causal inference.
problem Learning optimal policies in causal inference with limited data.
method RAG-based policy learning with vector search and nearest-neighbor matching.
result The methods bound the within-candidate choice regret and evaluate the one-step method directly as a policy.
CausalCOMRL improves RL task representations by integrating causal relationships, enhancing generalizability.
problem Spurious correlations in context-based offline meta-reinforcement learning.
method CausalCOMRL integrates causal representation learning to uncover and incorporate causal relationships among task components.
result CausalCOMRL achieves better performance on meta-reinforcement learning benchmarks.
This review explores causal decision-making to improve decision quality.
problem Effective decision-making requires understanding causal relationships.
method Causal structure learning, causal effect learning, and causal policy learning.
result Challenges in causal decision-making are identified and recent advances are discussed.
This paper improves causal inference using deep neural networks for low-dimensional covariates.
problem Improving causal inference with deep learning for high-dimensional covariates.
method Doubly robust off-policy learning with deep neural networks on low-dimensional manifolds.
result Nonasymptotic regret bounds for finite- and continuous-action scenarios, converging at a fast rate depending on intrinsic manifold dimension.
Causal reasoning has been an indispensable capability for humans and other intelligent animals to interact with the physical world. In this work, we propose to endow an artificial agent with the capability of causal reasoning for completing goal-directed tasks. We develop learning-based approaches to inducing causal kn…
Paper introduces effect-invariance for better policy generalization.
problem Adapting policies to unseen environments efficiently.
method Introduces effect-invariance, a relaxation of full invariance, and develops testing procedures to test e-invariance directly from data.
result Effect-invariance enables zero-shot and few-shot policy generalization without assuming a causal graph.
FOCUS improves offline RL by incorporating causal structure into world-models.
problem Learning effective policies from historical data without interaction.
method FOCUS proposes a practical algorithm that learns and leverages causal structure in offline RL.
result FOCUS outperforms plain model-based offline RL algorithms and other causal model-based RL algorithms.
DFPV improves PCL for confounded bandit policy evaluation.
problem Estimating causal effects in confounded settings with high-dimensional data.
method Deep feature proxy variable method (DFPV) for high-dimensional, nonlinear relationships.
result DFPV outperforms state-of-the-art methods on synthetic benchmarks and confounded bandit problems.
Deep learning analyzes healthcare provider actions and patient outcomes.
problem Understanding provider behavior in non-randomized healthcare settings.
method Deep causal behavioral policy learning (DC-BPL) using transformer architecture.
result Optimal provider policies identified for specific patient types.
Paper proposes a new method for demand forecasting in pricing contexts.
problem Demand forecasting in pricing contexts, especially in a profit optimal manner.
method Combines Double Machine Learning for causal inference and transformer-based forecasting models.
result Our method outperforms other forecasting methods in off-policy settings.
New method estimates and optimizes policy differences using orthogonal learning.
problem Offline reinforcement learning with safety concerns and cost limitations.
method Dynamic R-learner for estimating and optimizing Qπ(s,1)−Qπ(s,0), leveraging orthogonal estimation. result Consistent policy optimization with improved convergence rates.
Proposes a method to improve treatment policies in data-scarce clinical settings.
problem Improving treatment policies in data-scarce clinical settings with unobserved confounding.
method Uses a causal mechanism to model the underlying generative process and augments counterfactual trajectories with source domain priors.
result Significantly improves treatment policy performance in a simulated sepsis treatment task.
New method learns optimal policies in presence of unmeasured confounders.
problem Optimal policy learning with unobserved confounders.
method Causal-assisted policy learning methods using instrumental variables and negative controls.
result Policies are ildeO(n−1/2) quantile-optimal under mild coverage assumptions. New framework for dynamic causal graph modeling and effect estimation.
problem Dynamic changes in causal relationships over time.
method Score-based causal discovery with autoregressive model structure.
result Dynamic causal graph with time-varying causal relations.
The paper uses causal machine learning to optimize rework decisions in manufacturing.
problem Optimizing rework policies in manufacturing systems to balance yield improvement and rework costs.
method Proposes a causal model using double/debiased machine learning (DML) techniques to estimate conditional treatment effects and derive rework policies.
result Achieved a yield improvement of 2-3% during the color-conversion process of white LEDs.
Study introduces a new framework for policy learning without positivity assumption.
problem Learning optimal treatment assignment policies from observational data with constraints.
method Incremental propensity score policies and semiparametric efficiency theory.
result Validated framework's performance through numerical experiments.
Study uses causal machine learning to assess coupon campaign impact on retailer sales.
problem Assessing the causal effect of a coupon campaign on retailer sales.
method Causal machine learning algorithms, subgroup analysis, optimal policy learning.
result Only two coupon categories (drugstore and other food) have a significant positive impact on sales.
Proposes a method to estimate policy values in reinforcement learning with unmeasured confounders.
problem Estimating policy values in reinforcement learning with unmeasured confounders.
method Develops a two-way deconfounder algorithm using a neural tensor network to learn unmeasured confounders and system dynamics.
result Consistent policy value estimation through model-based estimator.
ALIAS uses RL to learn DAGs without acyclicity constraints.
problem Efficiently learning DAGs from observational data without acyclicity constraints.
method ALIAS employs RL to generate DAGs in a single step with optimal complexity, bypassing acyclicity constraints.
result ALIAS outperforms state-of-the-art methods in causal discovery.
This paper introduces an innovative Bayesian machine learning algorithm to draw interpretable inference on heterogeneous causal effects in the presence of imperfect compliance (e.g., under an irregular assignment mechanism). We show, through Monte Carlo simulations, that the proposed Bayesian Causal Forest with Instrum…
A-ICP selects experiments to learn causal effects efficiently.
problem Learning causal effects from observational data is difficult.
method Active learning framework based on Invariant Causal Prediction.
result Proposes intervention selection policies to reveal direct causes.
New approach to off-policy evaluation connects causal graph to policy effects.
problem Evaluating policies using observational data from different policies.
method Formalizes off-policy evaluation within a causal graph framework.
result Identifies specific causal estimands and highlights necessary experimental data.
Causal models bring many benefits to decision-making systems (or agents) by making them interpretable, sample-efficient, and robust to changes in the input distribution. However, spurious correlations can lead to wrong causal models and predictions. We consider the problem of inferring a causal model of a reinforcement…
A new causal deepset framework improves off-policy evaluation under complex interference.
problem Handling spatio-temporal interference in off-policy evaluation.
method Permutation invariance assumption and novel algorithms incorporating it.
result Significantly more precise estimations than existing methods.
The paper proposes a policy learning framework for interpretable personalization.
problem Effective personalization of goods and services to improve revenues and maintain competitive edge.
method Policy learning with linear decision boundaries using causal inference and Bayesian optimization.
result The learned policy improves net sales revenue by 88.2% and provides insights into important features.
This paper provides a link between causal inference and machine learning techniques - specifically, Classification and Regression Trees (CART) - in observational studies where the receipt of the treatment is not randomized, but the assignment to the treatment can be assumed to be randomized (irregular assignment mechan…
Counterfactual policy evaluation improves autonomous driving policies' generalization.
problem Learnt policies often fail to generalize and handle novel situations.
method Introduces counterfactual policy evaluation using counterfactual worlds.
result Significantly decreases collision-rate while maintaining high success-rate.
The paper tackles performative policy learning with strategic agents, improving scalability and generalizability.
problem Strategic agents adjust their features in response to a released policy, causing endogenous distribution shifts.
method Relaxing parametric assumptions, the paper uncovers a low-dimensional structure in distribution shifts and proposes a gradient-based policy optimization algorithm.
result The proposed algorithm achieves high sample efficiency and provides theoretical guarantees for convergence.
New framework for AI to learn causal models through experience.
problem Lack of guidance for variable choice and interventions in causal models for AI.
method Defines actions as state space transformations, introduces causal variables, and identifies interventions.
result Clarifies the concept of interventions and makes causal representation learning clearer.
Proposes a new method to estimate continuous treatment policies and match treatments effectively.
problem Current methods struggle with continuous treatment policies and complex matching.
method Formulates treatment effectiveness as a parametrizable model, using deep learning for optimization.
result Significant improvement in treatment effectiveness and matching efficiency.
We introduce an off-policy evaluation procedure for highlighting episodes where applying a reinforcement learned (RL) policy is likely to have produced a substantially different outcome than the observed policy. In particular, we introduce a class of structural causal models (SCMs) for generating counterfactual traject…
Reduces variance in noisy social outcomes to improve policy evaluation and optimization.
problem Improving access to opportunity through personalized treatment decisions.
method Data-driven dimensionality-reduction using reduced rank regression to denoise multiple outcomes.
result Improves estimation error in policy evaluation and optimization, including on real-world data.
Paper tackles transfer RL under unobserved context, developing methods to reduce bias.
problem Transfer RL with unobserved contextual information leading to biased models.
method Develops causal bounds on transition and reward functions using demonstrator's data.
result Proposes Q learning and UCB-Q learning algorithms that converge to true value function without bias.
While machine learning (ML) methods have received a lot of attention in recent years, these methods are primarily for prediction. Empirical researchers conducting policy evaluations are, on the other hand, pre-occupied with causal problems, trying to answer counterfactual questions: what would have happened in the abse…
Framework identifies causal factors of climate change using correlations and machine learning.
problem Understanding socioeconomic factors influencing carbon emissions and climate change.
method Three-step framework: correlation analysis, causal discovery, LLM interpretations.
result Adaptable solutions for data-driven policy-making and strategic decision-making.
New RL method learns policies from few data using causal models.
problem Limited interaction data and heterogeneous patient responses.
method Exploits structural causal models to model state dynamics and counterfactual reasoning.
result Counterfactual RL algorithms converge to optimal value function.
S-DIDML integrates structural DID with ML for causal inference in high-dimensional data.
problem Causal inference in high-dimensional observational panel data with confounding variables.
method Structural identification with high-dimensional estimation, Neyman orthogonality, cross-fitting, causal forests, semi-parametric models.
result Precision in identifying policy-sensitive groups and optimizing resource allocation.
The paper tackles counterfactual inference with multioutput deep kernels in high-dimensional settings.
problem Performing counterfactual inference with observational data in high-dimensional settings with multiple actions and outcomes.
method The paper presents a general class of counterfactual multi-task deep kernels models based on Structural Causal Models (SCM) and Gaussian Processes.
result The models estimate causal effects and learn policies efficiently, scaling well with high dimensions.
Causal Imitation Learning handles noisy measurements and distribution shifts.
problem Learning from noisy state observations and distributional shifts.
method Causal inference framework and adversarial RKHS learning.
result Improved robustness to distribution shifts compared to standard methods.
GO-CBED optimizes experiments for specific causal queries, improving efficiency.
problem Efficiently infer causal relationships with limited resources.
method Goal-oriented Bayesian framework that maximizes expected information gain on user-specified causal quantities.
result GO-CBED outperforms existing methods in various causal tasks, especially with limited budgets.
We present new methods to estimate causal effects retrospectively from micro data with the assistance of a machine learning ensemble. This approach overcomes two important limitations in conventional methods like regression modeling or matching: (i) ambiguity about the pertinent retrospective counterfactuals and (ii) p…
Develops adaptive framework for estimating survival effects with censoring.
problem Estimating causal effects in survival data with censoring.
method Derives semiparametric efficiency bound, proposes efficiency-optimal allocation policy, and develops Adaptive Survival Estimator (ASE).
result ASE achieves asymptotic normality via martingale central limit theorem and demonstrates efficiency gains over uniform randomization.