A new method removes biases in data integration by using surrogate control outcomes.
problem Data integration methods can be biased due to data-dependent processes.
method Post-integrated inference method using surrogate control outcomes to account for latent heterogeneity.
result The method provides consistent and efficient estimators under minimal assumptions and potential misspecifications.
We present a new machine learning approach to estimate personalized treatment effects in the classical potential outcomes framework with binary outcomes. To overcome the problem that both treatment and control outcomes for the same unit are required for supervised learning, we propose surrogate loss functions that inco…
Study uses surrogate data to improve treatment effect estimation with scarce outcome data.
problem Limited outcome data hinders estimating treatment effects.
method Uses abundant surrogate data to estimate treatment effects without stringent assumptions.
result Improves precision of treatment effect estimation.
Estimating the long-term effects of treatments is of interest in many fields. A common challenge in estimating such treatment effects is that long-term outcomes are unobserved in the time frame needed to make policy decisions. One approach to overcome this missing data problem is to analyze treatments effects on an int…
New method uses AI predictions as cheaper alternatives to expensive outcomes.
problem Using expensive outcomes for statistical inference.
method Recalibrated prediction-powered inference using machine learning techniques.
result Significant gains in effective sample size over existing PPI proposals.
The paper explores trading off consistency and dimensionality in convex surrogates for multiclass classification.
problem Designing consistent surrogate losses for multiclass classification with high-dimensional outcomes.
method Investigates embedding outcomes into convex polytopes and examining consistency under low-noise assumptions.
result Consistency can be achieved with less than n−1 dimensions, but hallucination occurs for some distributions. The paper proposes a method to estimate treatment effects using surrogates when primary outcomes are missing.
problem Missing primary outcomes in causal inference applications can lead to biased estimates.
method Doubly robust method that uses both labeled and unlabeled data, incorporating surrogates.
result The proposed estimator is asymptotically normal and has improved variance compared to methods using only labeled data.
New method uses surrogate outcomes and single-record data to improve suicide risk modeling.
problem Lack of historical information in single-record patients hinders modeling rare medical events.
method Hybrid framework combining supervised and unsupervised learning to integrate concurrent and single-record data.
result Single-record data and concurrent diagnoses provide valuable information for improving suicide risk modeling.
Develops a SAS approach for high-dimensional risk prediction using unlabeled data.
problem Challenges in risk modeling with EHR data due to lack of direct disease outcomes and high dimensionality.
method Surrogate Assisted Semi-supervised Learning (SAS) approach leveraging unlabeled and labeled data.
result Valid inference for predicted risk even when underlying model is dense and mis-specified.
The paper proposes a method to assess surrogate heterogeneity in non-randomized data.
problem Lack of methods to evaluate surrogate heterogeneity in non-randomized data.
method Proposes a framework using meta-learners to assess surrogate heterogeneity in real-world data.
result Identifies individuals for whom the surrogate is a valid replacement of the primary outcome.
Fuses ITRs for primary and secondary outcomes to minimize harm.
problem Learn an ITR maximizing primary outcome while minimizing harm to secondary outcomes.
method Introduces fusion penalty to encourage similar recommendations for different outcomes. Two algorithms estimate the ITR using surrogate loss functions.
result Agreement rate between primary and secondary optimal ITRs converges faster than ignoring secondary outcomes.
Develops a numerical algorithm for stochastic impulse control using regression surrogates.
problem Optimal impulse control in stochastic processes.
method Generates statistical surrogates for continuation and intervention functions, recursively trained over simulated state trajectories.
result Demonstrates flexibility and extensibility of the numerical scheme through case studies.
This paper investigates the control of an ML component within the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) devoted to black-box optimization. The known CMA-ES weakness is its sample complexity, the number of evaluations of the objective function needed to approximate the global optimum. This weakness is…
Efficient RL method uses SSL for small labeled data in health outcomes.
problem Lack of precise health outcome data for reinforcement learning.
method Semi-supervised learning approach for Q-learning and value estimation.
result Method efficiently estimates Q-function and value function, robust to mis-specification.
ConfEviSurrogate improves surrogate model accuracy and uncertainty quantification.
problem Uncertainty in surrogate models hinders reliable analysis.
method Introduces ConfEviSurrogate, a novel model that learns evidential distributions, separates uncertainty sources, and provides reliable prediction intervals.
result Demonstrates accurate predictions and robust uncertainty estimates in various simulations.
This paper proposes a method for estimating the effect of a policy intervention on an outcome over time. We train recurrent neural networks (RNNs) on the history of control unit outcomes to learn a useful representation for predicting future outcomes. The learned representation of control units is then applied to the t…
Bayesian framework for policy learning in decision problems.
problem Maximizing expected welfare in decision-making problems.
method Loss-based Bayesian updating and squared-loss surrogate for welfare maximization.
result General Bayes posterior over decision rules with Gaussian pseudo-likelihood interpretation.
In a seminal paper Abadie, Diamond, and Hainmueller [2010] (ADH), see also Abadie and Gardeazabal [2003], Abadie et al. [2014], develop the synthetic control procedure for estimating the effect of a treatment, in the presence of a single treated unit and a number of control units, with pre-treatment outcomes observed f…
A new method models continuous-time counterfactual outcomes using neural controlled differential equations.
problem Estimating personalized healthcare outcomes over irregularly sampled data.
method Interpreting data as samples from a continuous-time process, modeling latent trajectory using controlled differential equations, and using adversarial training for time-dependent confounding.
result TE-CDE consistently outperforms existing approaches in irregularly sampled scenarios.
We present a framework for automatically structuring and training fast, approximate, deep neural surrogates of stochastic simulators. Unlike traditional approaches to surrogate modeling, our surrogates retain the interpretable structure and control flow of the reference simulator. Our surrogates target stochastic simul…
A new meta-algorithm for estimating the conditional average treatment effects is proposed in the paper. The main idea underlying the algorithm is to consider a new dataset consisting of feature vectors produced by means of concatenation of examples from control and treatment groups, which are close to each other. Outco…
Derivative-informed models improve financial surrogates for accurate hedging and risk management.
problem Developing fast surrogate models for financial derivatives and risk quantities.
method Derivative-informed operator learning framework combining neural operators, random features, and tangent sensitivity equations.
result The framework reduces hedging and risk errors by 40-76% compared to standard surrogates.
This paper develops efficient surrogate models for optimization of complex dynamical systems.
problem Computational expense in solving complex dynamical systems through numerical simulation.
method Combination of proper orthogonal decomposition and radial basis functions for constructing low-dimensional surrogate models.
result Surrogate models reduce computational time for optimization problems while maintaining accuracy.
Study risk-constrained Kelly optimization for mutually exclusive outcomes, proving support invariance and developing a structured algorithm.
problem Risk-constrained Kelly optimization for mutually exclusive outcomes with explicit state prices.
method Analyzes the finite mutually exclusive outcome version of risk-constrained Kelly optimization with explicit state prices, proving support invariance and developing a structured algorithm.
result Support is invariant across CRRA parameter and drawdown-surrogate parameter in the overround regime.
Optimizes experimental design using synthetic controls for better outcomes.
problem Estimating average treatment effects in studies with pre-treatment data.
method Mixed-integer programming for selecting treated and control units and weights.
result Improves mean squared error and statistical power compared to simple alternatives.
Kernel methods identify treatment effects with unobserved confounding using negative controls.
problem Learning causal relationships with unmeasured confounding.
method Kernel ridge regression algorithms for nonparametric treatment effects.
result Uniform consistency and finite sample rates of convergence proved.
New method tackles model uncertainty in stochastic control using Bayesian nonparametrics.
problem Model uncertainty in stochastic control problems.
method Nonparametric Bayesian approach with Dirichlet process for unknown distributions, online learning, and Gaussian process surrogates.
result Demonstrates financial advantages of nonparametric Bayesian over parametric methods.
We develop DTs for PDE models using KL-NN and TL, analyzing TL's moment equations and one-shot learning for exactness.
problem Creating accurate digital twins for systems governed by PDEs under changing conditions.
method We use KL-NN surrogate models and transfer learning to construct DTs, analyzing the moment equations and proposing one-shot and few-shot learning methods.
result For linear PDEs, one-shot TL is exact; for nonlinear PDEs, some parameters can be transferred with minimal error.
In the absence of unobserved confounders, matching and weighting methods are widely used to estimate causal quantities including the Average Treatment Effect on the Treated (ATT). Unfortunately, these methods do not necessarily achieve their goal of making the multivariate distribution of covariates for the control gro…
TSC improves causal effect estimation in panel data.
problem Estimating causal effects in panel data with a single treated unit.
method Targeted synthetic control method that refines initial weights through a one-dimensional targeted update.
result TSC consistently improves estimation accuracy over state-of-the-art SCM baselines.
New method improves model reconstruction using counterfactuals and polytope theory.
problem Reconstructing models with minimal input changes and avoiding decision boundary shifts.
method Using polytope theory to derive loss functions that treat counterfactuals differently from ordinary instances.
result Improves fidelity between target and surrogate model predictions on multiple datasets.
Differentiable simulations control molecular Hamiltonians for desired outcomes.
problem Control and learning of molecular Hamiltonians for desired outcomes.
method Differentiable simulations to differentiate Hamiltonians with respect to target observables.
result Control and learning of molecular Hamiltonians for desired outcomes.
The paper targets optimal interventions for long-term outcomes using imputed data and policy learning.
problem Maximizing long-term outcomes observed only in the future.
method Imputing missing long-term outcomes and using a doubly-robust approach for policy evaluation and optimization.
result The approach outperforms simple short-term proxies and achieves significant revenue impact over three years.
This study evaluates subgroup analysis methods for time-to-event outcomes in randomized controlled trials.
problem Identifying subgroups of good responders in non-significant randomized controlled trials.
method Evaluation of several subgroup analysis algorithms for time-to-event outcomes using synthetic and semi-synthetic data.
result Provides a new synthetic and semi-synthetic data generation process and an open-source Python package for benchmarking.
A new method TNW-CATE estimates treatment effects using neural networks.
problem Estimating heterogeneous treatment effects with limited controls and many treatments.
method Trainable Nadaraya-Watson regression with shared parameters neural network.
result TNW-CATE outperforms traditional methods in various simulation experiments.
New method uses geometric mean to avoid non-collapsibility in case-control studies.
problem Non-collapsibility of odds ratio under outcome-dependent sampling.
method Proposes geometric mean aggregation to avoid non-collapsibility and provides estimation and inference methods.
result Geometric odds ratio is collapsible under outcome-dependent sampling.
Machine learning reduces variance in online experiment results.
problem Reducing variance in randomized controlled trials.
method Machine learning regression-adjusted treatment effect estimator (MLRATE).
result MLRATE reduces estimator variance by over 70% in A/A tests.
Surrogate-based analysis of interactions via local effect smooths
problem Detecting and characterizing feature interactions in machine learning models
method Surrogate-based analysis using generalized additive models
result Empirical validation of effectiveness for pairwise interactions
A new beta-VAE based regression model accelerates oilfield optimization studies.
problem Computational expense of full-physics reservoir simulations.
method beta-VAE for interpretable latent space representation, probabilistic dense layers for uncertainty quantification.
result Interpretable latent representation and quantified uncertainty for optimization decisions.
Conformal Candidate Certification advances offline MBO by certifying candidate designs with statistical guarantees.
problem Offline model-based optimization
method Conformal Candidate Certification (CCC)
result CCC certifies 16.7% of an aggressive proposal pool with 0.990 empirical coverage at nominal 0.90.
The paper introduces a new algorithm for fair decision-making in outcome control tasks.
problem Fair and equitable automated decision-making in outcome control tasks.
method Causal analysis and optimization to ensure fairness in decision-making.
result Developed an algorithm for maximizing Y while ensuring causal fairness. Firms delay write-downs for adverse macroeconomic and industry outcomes but not for firm-specific issues.
problem Timeliness of write-downs for adverse macroeconomic and industry outcomes versus firm-specific issues.
method Comparative analysis of write-downs driven by macroeconomic and industry outcomes versus firm-specific outcomes.
result Firms delay write-downs for adverse macroeconomic and industry outcomes but not for firm-specific issues.
Synthetic control method improves policy evaluation in high-dimensional settings.
problem Evaluating the impact of new policies in large-scale applications.
method Two-phase approach: nearest neighbor matching followed by supervised learning.
result The method successfully improves estimate accuracy in large-scale experiments.
New method reduces variance in Bayesian inverse problems.
problem High variance in Monte Carlo estimates for inverse problems.
method Conditional neural control variates based on Stein's identity.
result Substantial variance reduction across different inverse problems.
Two-stage TMLE reduces bias and improves efficiency in CRTs.
problem Differential outcome measurement and imbalance in baseline predictors in CRTs.
method Two-stage targeted minimum loss-based estimator (TMLE) to adjust for baseline covariates.
result Our approach nearly eliminates bias due to differential outcome measurement.
The paper proposes a neural network method to estimate treatment effects by balancing treated and control distributions.
problem Estimating individual and average treatment effects from observational data.
method Balance regularization of multi-head neural network architectures to reduce confounding effects.
result The approach reduces bias-variance trade-off and improves treatment effect estimation.
This paper describes Plumbing for Optimization with Asynchronous Parallelism (POAP) and the Python Surrogate Optimization Toolbox (pySOT). POAP is an event-driven framework for building and combining asynchronous optimization strategies, designed for global optimization of expensive functions where concurrent function …
Tests whether a treatment's effect is fully mediated by observed outcomes and identifies causal mechanisms.
problem Understanding how a treatment affects an outcome through intermediate variables.
method Proposes a test to evaluate full mediation and causal mechanism identification, extending to non-randomly assigned treatments.
result A conditionally random treatment is conditionally independent of the outcome given mediators and covariates if full mediation and causal mechanism identification hold.