Existing imitation learning approaches often require that the complete demonstration data, including sequences of actions and states, are available. In this paper, we consider a more realistic and difficult scenario where a reinforcement learning agent only has access to the state sequences of an expert, while the expe…
Study tackles OPE in confounded settings, estimating policy value from proxies.
problem Difficulty in OPE due to unobserved confounders in infinite-horizon RL.
method Two-stage approach: estimating stationary distribution ratios and combining optimal balancing.
result Policy value can be identified from off-policy data with proxies and latent variable model.
This work highlights problems with off-policy estimation in recommender systems due to unobserved confounders.
problem Evaluation of recommender systems under unobserved confounders.
method Policy-based estimators and characterisation of statistical bias due to confounding.
result Naive propensity estimation under confounding leads to severely biased metric estimates.
New method for robust policy evaluation in offline reinforcement learning with sequentially exogenous unobserved confounders.
problem Offline reinforcement learning in domains with unobserved confounders.
method Orthogonalized robust fitted-Q-iteration with closed-form solutions and bias-correction.
result Effective in simulations and real-world data, improving robustness and computational ease.
New algorithm learns optimal decisions from imperfectly observed contexts.
problem Learning optimal decisions in bandits with unobserved contexts.
method Posterior sampling algorithm for imperfectly observed contexts.
result Efficient learning from noisy imperfect observations.
New method predicts unobserved interactions between sets of elements.
problem Limited access to interactions between sets of elements.
method Generalized Synthetic Interventions (GSI) estimator.
result GSI estimator outperforms existing methods on synthetic and real data.
Develops a method to estimate policy values robustly in the presence of confounding variables.
problem Infinite-horizon reinforcement learning with unobserved confounding variables makes policy evaluation unidentifiable.
method Robust approach estimating sharp bounds on policy value using optimization over state-occupancy ratios and sensitivity model.
result Proves convergence to sharp bounds as more confounded data is collected.
Paper identifies unobserved variables from observable data.
problem Missing variables in empirical studies.
method Function mapping from observables to unobservables based on joint distribution.
result Uniqueness of latent values in each observation.
New methods use vector search and nearest-neighbor matching for policy learning in causal inference.
problem Learning optimal policies in causal inference with limited data.
method RAG-based policy learning with vector search and nearest-neighbor matching.
result The methods bound the within-candidate choice regret and evaluate the one-step method directly as a policy.
We consider an investor faced with the utility maximization problem in which the risky asset price process has pure-jump dynamics affected by an unobservable continuous-time finite-state Markov chain, the intensity of which can also be controlled by actions of the investor. Using the classical filtering theory, we redu…
KRCD detects unobserved confounders in nonlinear observational data.
problem Detecting unobserved confounders in nonlinear observational studies.
method Kernel Regression Confounder Detection (KRCD) using reproducing kernel Hilbert spaces.
result KRCD outperforms existing methods and achieves superior computational efficiency.
New method for finding optimal treatment regimes in medical settings with time-varying unobserved factors.
problem Finding optimal treatment regimes in medical settings with time-varying unobserved factors.
method Extend Dynamic Treatment Regimes (DTRs) to Ambiguous Dynamic Treatment Regimes (ADTRs), connect to Ambiguous Partially Observable Mark Decision Processes (APOMDPs), and develop Reinforcement Learning methods.
result Established theoretical results for learning methods, including consistency and asymptotic normality.
Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actions in the demonstrations to be fully available, which is hard to ensure in real applications. Though algorithms for learning with unobservab…
Valid causal inference with unobserved confounding in high-dimensional settings.
problem Estimating causal effects with unobserved confounders in high-dimensional data.
method Proposes methods to estimate causal effects with valid confidence intervals in the presence of unobserved confounders and high-dimensional nuisance models.
result Valid semiparametric inference can be obtained with unobserved confounding, and uncertainty intervals are proposed.
A new method uses randomized trials to estimate the strength of unobserved confounding.
problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.
New method estimates treatment effects over time with unobserved confounders.
problem Estimating treatment effects from observational data with unobserved confounders.
method Sequential Deconfounder using Gaussian process latent variable model.
result Unbiased estimates of individualized treatment responses over time.
Study online learning in unknown Markov games with sublinear regret.
problem Online learning in unknown Markov games with unobservable opponents.
method Introduced an algorithm achieving sublinear regret against the minimax value.
result First sublinear regret bound for unknown Markov games, independent of action spaces size.
New method scores DAGs by identifying unobserved confounding.
problem Unobserved confounding complicates causal discovery.
method Score-based causal discovery algorithm that accounts for unobserved confounding.
result Sparse linear Gaussian DAGs can be recovered from observed data.
New method recovers predictions from unobservable source subpopulation in binary classification.
problem Challenging binary classification with unobservable subpopulation in source domain.
method Distribution matching method to estimate subpopulation proportions, rigorous derivation of prediction models.
result Our method outperforms naive benchmarks in synthetic and real-world datasets.
Anomaly detection method tackles hidden adversary actions in real-world scenarios.
problem Detecting hidden adversary actions in real-world scenarios where the defender cannot perfectly observe the attacker's actions.
method Extends existing anomaly detection models to handle continuous action spaces and game-theoretic framework. Proposes two algorithms: direct extension and learning-based approach.
result Learning-based approach produces less exploitable strategies and is scalable to higher dimensions.
New method estimates policy performance under unobserved confounding.
problem Estimating policy performance when decisions depend on unobserved variables.
method Developed worst-case bounds for robust OPE under unobserved confounding.
result Efficient procedure for computing worst-case bounds, proving statistical consistency.
Proposes a new method for algorithmic recourse in confounded settings.
problem Provides actionable recommendations for individuals affected by automated decisions.
method Relaxes assumptions of no hidden confounding and additive noise, requiring only causal graph and confounding structure.
result Bounds the expected counterfactual effect of recourse actions, ensuring favourable outcomes in expectation.
CDVAE estimates treatment effects over time by accounting for unobserved variables.
problem Estimating treatment effects over time in the presence of unobserved confounders.
method Causal Dynamic Variational Autoencoder (CDVAE) that addresses unconfoundedness and unobserved heterogeneity.
result CDVAE outperforms existing methods in estimating Conditional Average Treatment Effects (CATEs).
Paper tackles unobserved confounding in human-AI collaborations.
problem Unobserved confounding undermines human-AI collaboration effectiveness.
method Combines sensitivity analysis from causal inference with AI-driven statistical modeling.
result Enhances robustness and reliability of collaborative outcomes.
Proposes ρ ρ ρ -GNF for sensitivity analysis of unobserved confounding.
problem Sensitivity analysis of unobserved confounding in observational studies.
method Copulas and normalizing flows to estimate average causal effect (ACE) as a function of unobserved confounding strength.
result Develops ρ c u r v e ρ_{curve} ρ c u r v e to provide bounds for ACE and identify confounding strength required to nullify ACE. Paper adapts DML for panel data, addressing unobserved heterogeneity.
problem Estimating causal effects with panel data and unobserved heterogeneity.
method Adapting double/debiased machine learning (DML) for panel data with predictive models based on correlated random effects.
result Predictive models based on correlated random effects within DML lead to accurate coefficient estimates.
New method removes hidden confounders for unbiased treatment effect estimation.
problem Bias in treatment effect estimation due to unobserved confounders.
method Proposes a new debiased estimation approach via SVD to handle heterogeneous confounding.
result Established rate of convergence for the estimator under different noise conditions.
The paper tackles robust domain generalization by accounting for unobserved confounders.
problem Learning robust, generalizable models from multiple datasets in the presence of unobserved confounders.
method Defines a new invariance property for causal solutions, connects it to distributionally robust optimization, and incorporates regularization to encourage partial equality of error derivatives.
result Demonstrates the empirical effectiveness of the approach on healthcare data from various modalities.
Improved method for unbiased causal discovery in presence of unobserved confounding.
problem Unbiased data synthesis for causal discovery algorithms in the presence of unobserved confounding.
method Explicit block-hierarchical ancestral sampling to address limitations of implicit parameterization.
result Our approach fully covers the space of causal models, including those generated by implicit parameterization.
New framework for estimating treatment effects in observational studies.
problem Estimating average treatment effects in the presence of unobserved confounders.
method Distributionally robust optimization, sensitivity models.
result Sharp bounds on average treatment effects under distributional assumptions.
GUM tackles MARL by avoiding overestimation through state-marginal restriction.
problem Overestimation of values in large joint state-action spaces.
method Greedy UnMixing through state-marginal restriction and unmixing.
result Superior performance compared to existing Q-learning and general MARL algorithms.
New method identifies causal effects with categorical unobserved confounders.
problem Estimating causal effects in the presence of unobserved confounders.
method Mixture learning and tensor decomposition for consistent estimation.
result Causal effects are identifiable with categorical unobserved confounders under suitable conditions.
The paper shows how to audit fairness in decisions with hidden risk factors.
problem Estimating fairness in decisions influenced by hidden, unobservable risk factors.
method Derives unbiased estimates of risk using historical data and audits existing decision-making systems.
result One can compute meaningful bounds on treatment rates for high-risk individuals, even with hidden confounders.
Develops Austen plots for assessing bias from unobserved confounding in observational studies.
problem Bias in causal estimates due to unobserved confounding.
method Formalizes confounding strength, uses Austen plots to visualize and quantify bias.
result Allows domain experts to assess the plausibility of strong confounders.
Causal discovery predicts unobserved joint statistics from observed data.
problem Inferring properties of unobserved joint distributions from observed data.
method Infer causal models from observed data to predict statistical properties of unobserved sets.
result Sparse causal graphs can be more useful than dense ones in predicting unobserved joint distributions.
A new algorithm CAP learns optimal policies from observational data with confounding bias and missing observations.
problem Offline contextual bandit with confounding bias and missing observations.
method CAP policy learning, forming reward function as solution of integral equation system, building confidence set, and greedily taking action with pessimism.
result Developed an upper bound to the suboptimality of CAP for the offline contextual bandit problem.
Study of repeated games with unobserved agent rewards using MAB framework.
problem Designing policies for principals in repeated principal-agent games with unobservable agent rewards.
method Developed a policy achieving low regret (square-root regret up to a log factor) for perfect-knowledge agents.
result Constructed an estimator for agent's expected reward and designed a policy achieving low regret.
New bounds assess policy evaluation under unobserved confounders, showing model-based methods are more effective.
problem Policy evaluation under unobserved confounders in uncertain causal environments.
method Developed worst-case bounds for sensitivity to unobserved confounders, demonstrating model-based methods are more effective.
result Model-based approaches with robust MDPs provide sharper lower bounds for policy evaluation.
In many applications of network analysis, it is important to distinguish between observed and unobserved factors affecting network structure. To this end, we develop spectral estimators for both unobserved blocks and the effect of covariates in stochastic blockmodels. On the theoretical side, we establish asymptotic no…
PRL improves off-policy evaluation in partially observed MDPs.
problem Confounding and bias in offline reinforcement learning with unobserved state factors.
method Extends proximal causal inference to POMDPs, identifying and estimating target policy value.
result Semiparametrically efficient estimators for PRL in partially observed MDPs.
Proposes a method to assess unobserved confounding effects in causal inference.
problem Assessing unobserved confounding in causal inference studies.
method Copula-based normalizing flows with sensitivity parameter ρ ρ ρ . result Estimates average causal effect (ACE) as a function of unobserved confounding strength.
A new algorithm uses IVs to learn optimal policies from observational data.
problem Learning optimal policies from unobserved variable confounded data.
method IV-aided Value Iteration (IVVI) algorithm based on conditional moment restrictions.
result First provably efficient algorithm for instrument-aided offline RL.
Intact-VAE estimates treatment effects with latent confounders.
problem Estimating treatment effects under unobserved confounding.
method Intact-VAE, a VAE variant, models latent confounders to identify treatment effects.
result Intact-VAE is a consistent estimator of treatment effects under certain settings.
This work restricts hidden cardinality in causal models to infer causal relations.
problem Causal relations between variables with a common unobserved cause cannot be directly inferred.
method Derive inequality constraints from d-separation in causal models with known cardinalities of unobserved variables.
result Inference of causal relations is possible with additional assumptions about cardinalities.
We develop a cross-sectional research design to identify causal effects in the presence of unobservable heterogeneity without instruments. When units are dense in physical space, it may be sufficient to regress the "spatial first differences" (SFD) of the outcome on the treatment and omit all covariates. The identifyin…
We study Exo-MDPs to reduce sample complexity in reinforcement learning.
problem Reducing sample complexity in reinforcement learning for structured MDPs.
method Introducing Exo-MDPs and proving structural equivalence to linear mixture MDPs, establishing regret bounds.
result Proved O ( H 3 / 2 d K ) O(H^{3/2}d\sqrt{K}) O ( H 3/2 d K ) regret bound for Exo-MDPs, matching lower bounds. Method learns causal effects from multiple interventions in presence of unobserved confounders.
problem Disentangling causal effects from sets of interventions in the presence of unobserved confounders.
method Non-linear structural causal models with additive, multivariate Gaussian noise; algorithm that learns causal model parameters by pooling data from different regimes and maximizing combined likelihood.
result Identification proofs demonstrate that causal effects of single interventions can be learned from sets of interventions, even with unobserved confounders.
We identify causal models with unobserved confounding using bijective generation mechanisms.
problem Identifying causal relationships with unobserved confounders.
method Establish counterfactual identifiability for BGMs and propose a learning method.
result Learned BGMs enable efficient counterfactual estimation.