Paper addresses fairness issues in error-prone outcomes.
problem Fairness in error-prone outcomes.
method Combining fair ML methods and measurement models.
result Using a latent variable model removes detected unfairness.
Paper simplifies balancing weights by relaxing outcome assumptions.
problem Estimating missing outcomes in a target population.
method Relaxes outcome assumptions to simplify balancing weights.
result Balancing weights can be simplified with convex loss and minimum worst-case bias.
DCMA uses generative models to analyze treatment effects on entire outcome distributions.
problem Traditional mediation analysis focuses on summary contrasts, missing complex distributional changes.
method DCMA learns conditional generative models for mediators and outcome, reconstructing interventional distributions via Monte Carlo simulation.
result DCMA captures both summary effects and rich distributional contrasts like energy distance and Wasserstein distance.
DCMA uses generative models to analyze complex treatment effects on outcome distributions.
problem Analyzing complex and nonlinear causal mechanisms through outcome-level summary contrasts.
method Generative learning framework for identifying and estimating treatment effects on entire outcome distributions.
result Reconstructs interventional outcome distributions via Monte Carlo forward simulation, capturing both summary and distributional contrasts.
The paper discusses selecting predictive models for causal inference, highlighting the challenges and proposing a solution.
problem Selecting the best predictive models for causal inference from a variety of machine learning models.
method The paper proposes using Rext−risk, flexible estimators, and splitting data to compute risks for model selection. result The proposed method controls both outcome errors for treated and non-treated individuals, addressing the issue of model selection for causal inference.
Ex ante forecast outcomes should be interpreted as counterfactuals (potential histories), with errors as the spread between outcomes. Reapplying measurements of uncertainty about the estimation errors of the estimation errors of an estimation leads to branching counterfactuals. Such recursions of epistemic uncertainty …
Optimizes data labeling for causal effect estimation with missing outcomes.
problem Estimating causal effects with missing outcome data and budget constraints.
method Optimizes batch sampling probability to minimize variance of causal inference estimator.
result Achieves lower mean-squared error with fewer labeled data points.
A method corrects bias in estimating a high-dimensional classification rule using auxiliary outcomes.
problem Bias in estimating a high-dimensional classification rule using only one outcome.
method Robust transfer learning approach combining MTL and calibration steps.
result Final estimator achieves lower error than using only the target outcome.
We consider the problem of how to assign treatment in a randomized experiment, in which the correlation among the outcomes is informed by a network available pre-intervention. Working within the potential outcome causal framework, we develop a class of models that posit such a correlation structure among the outcomes. …
Proposes a method to improve CATE estimation by imputing missing potential outcomes.
problem Statistical discrepancy between distinct treatment groups in CATE estimation.
method Contrastive learning approach to reliably impute missing potential outcomes for a subset of individuals.
result Improves the accuracy and robustness of CATE estimation models.
New method improves counterfactual distribution learning for high-dimensional outcomes.
problem Counterfactual distribution learning for high-dimensional outcomes with concentrated structure.
method Geometry-adaptive diffusion-guided smoothing estimators combining causal nuisance adjustment and local outcome geometry.
result Geometry-adaptive methods show steeper error decay in semi-synthetic experiments.
Reduces variance in noisy social outcomes to improve policy evaluation and optimization.
problem Improving access to opportunity through personalized treatment decisions.
method Data-driven dimensionality-reduction using reduced rank regression to denoise multiple outcomes.
result Improves estimation error in policy evaluation and optimization, including on real-world data.
Paper estimates regression errors without known ground truth.
problem Estimating regression errors when ground truth is unknown.
method Developed a framework to estimate generalization error.
result Framework robustly detects concept drift in real-world datasets.
Paper proposes a method to model health outcomes using varying-coefficients and KNN-based LASSO.
problem Modeling health outcomes like BMI and cholesterol levels with varying age effects.
method Varying-coefficients regional quantile regression via KNN fused LASSO, with ADMM algorithm.
result Efficacy in capturing complex age-dependent associations between health outcomes and risk factors.
Study clarifies variance of stratification estimators for causal effects.
problem Estimating average causal effects with discrete covariates.
method Combines insights from potential outcomes, causal diagrams, and structural models.
result Derives expressions for the variance of stratification estimators.
The study uses transfer learning to compare surgical outcomes across racial/ethnic subgroups.
problem Difficulty in comparing surgical outcomes due to racial/ethnic and geographic differences.
method Causal inference framework and transfer learning to incorporate data from multiple populations.
result Racial and ethnic differences in surgical outcomes are found, with non-Hispanic Black patients experiencing wide variability.
It is well known that the distribution of returns from various financial instruments are leptokurtic, meaning that the distributions have "fatter tails" than a Normal distribution, and have skew toward zero. This paper presents a graceful micro-level explanation for such fat-tailed outcomes, using agents whose private …
A new method removes biases in data integration by using surrogate control outcomes.
problem Data integration methods can be biased due to data-dependent processes.
method Post-integrated inference method using surrogate control outcomes to account for latent heterogeneity.
result The method provides consistent and efficient estimators under minimal assumptions and potential misspecifications.
Efficiently generates models resistant to falsification.
problem Creating models that cannot be disproven by tests.
method Exploits connections between high-dimensional multicalibration and expected variational inequality problems to develop an efficient algorithm.
result First to efficiently produce online outcome indistinguishable generative models resistant to infinite classes of tests.
Cross-balancing improves causal inference by balancing features with outcome data.
problem Balancing features for valid causal inference when outcome data is available.
method Cross-balancing using sample splitting to separate feature construction and weight estimation errors.
result Cross-balancing produces consistent, asymptotically normal, and efficient estimators under mild conditions.
Improves matrix completion by exploiting biased observation patterns.
problem Matrix completion with biased observation patterns.
method Mask Nearest Neighbor (MNN) algorithm: two-stage process.
result MNN achieves competitive performance with 28x smaller mean squared error.
DynForest R package predicts outcomes with time-dependent predictors.
problem Handling time-dependent predictors in random forest models.
method Random forests with time-dependent predictors summarized using flexible linear mixed models.
result DynForest can predict continuous, categorical, and survival outcomes.
Estimates and tests treatment effects on entire outcome distributions.
problem Treatment effects on entire outcome distributions, not just averages.
method Proposes a novel estimand and doubly robust estimator, develops a test.
result First test with provably valid type 1 error guarantees in this setting.
New method robust to random distributional shifts in prediction.
problem Random distributional shifts in real-world settings.
method Hybrid approach combining long-term and proxy outcomes.
result Hybrid approach yields lower mean-squared error than current methods.
The paper shows how demographic data can lead to biased predictions, proposing 'Affirmative Information' as a solution.
problem Bias in predictions due to demographic data.
method Characterization of error types and conditions leading to disparate impact.
result Demographic variables in data can lead to biased predictions, with higher average outcomes receiving higher false positive rates.
New algorithms adaptively calibrate predictions in non-stationary environments, matching optimal rates.
problem Designing online prediction algorithms that adapt to varying levels of non-stationarity.
method Epoch-based scheduling and non-uniform partitioning of the prediction space.
result Achieves adaptive calibration guarantees under multiple measures with optimal rates.
Study fairness in intervention to maximize outcomes.
problem Fairness in intervention on a given node.
method Counterfactual estimation with partial causal model knowledge.
result Theoretical guarantees on error probability and effectiveness of algorithm.
We compare various extensions of the Bradley-Terry model and a hierarchical Poisson log-linear model in terms of their performance in predicting the outcome of soccer matches (win, draw, or loss). The parameters of the Bradley-Terry extensions are estimated by maximizing the log-likelihood, or an appropriately penalize…
New method for estimating out-of-sample R² from gene expression data.
problem Lack of a well-defined and unbiased estimator for out-of-sample R².
method Explicitly defined out-of-sample R², provided an unbiased estimator, and calculated standard error.
result Demonstrated improved model comparison for gene expression phenotypes.
The Artificial Prediction Market is a recent machine learning technique for multi-class classification, inspired from the financial markets. It involves a number of trained market participants that bet on the possible outcomes and are rewarded if they predict correctly. This paper generalizes the scope of the Artificia…
Machine learning can help personalized decision support by learning models to predict individual treatment effects (ITE). This work studies the reliability of prediction-based decision-making in a task of deciding which action a to take for a target unit after observing its covariates x~ and predicted outcom…
Bayesian method reduces misclassification errors in ranking Pareto-optimal solutions.
problem Identifying true Pareto-optimal solutions in noisy multiobjective optimization.
method Sequential allocation of extra samples using stochastic kriging to build predictive distributions.
result The proposed method outperforms existing algorithms in reducing misclassification errors.
New framework for network regression models accounting for community structure.
problem Inaccurate modeling of residual dependencies in network regression models.
method Modeling errors as community-based and exploiting exchangeability properties.
result Parsimonious standard errors for regression parameters.
A method for safe online classification reduces test costs while maintaining low error rates.
problem Sequential testing for binary disease outcomes with unknown logistic model parameters.
method Joint estimation of logistic parameter and feature distribution with a conservative threshold.
result Achieves target error with high probability and requires minimal excess tests.
DSL estimates heterogeneous treatment effects over time in survival settings.
problem Complicated by right censoring and time-varying treatment effects.
method Deep survival learner (DSL) for estimating CATEs over a clinically relevant time spectrum.
result DSL reveals heterogeneity in perioperative chemotherapy effects over time.
New method minimizes decision errors in large treatment spaces.
problem Improving decision-making in large treatment spaces with biased observational data.
method Loss minimizes classification error of actions in large action space.
result Proves improved decision-making performance in large combinatorial action spaces.
The volume of stroke lesion is the gold standard for predicting the clinical outcome of stroke patients. However, the presence of stroke lesion may cause neural disruptions to other brain regions, and these potentially damaged regions may affect the clinical outcome of stroke patients. In this paper, we introduce the t…
Machine learning improves kidney transplant outcomes prediction.
problem Improving prediction of kidney transplant success.
method Random forest machine learning model trained on kidney donor risk index data.
result Random forest predicted 2,148 more successful transplants than the risk index.
New estimator for causal effects in large datasets.
problem Unobserved confounding in large-scale data.
method Doubly robust estimator combining imputation, IPW, and cross-fitting.
result Error converges to Gaussian distribution at parametric rate.
The paper studies causal effects of multiple treatments in healthcare databases with rare outcomes.
problem Estimating causal effects of multiple treatments in healthcare databases with rare outcomes.
method The paper designs three sets of simulations and compares the operating characteristics of three types of methods: Bayesian Additive Regression Trees (BART), regression adjustment on multivariate spline of generalized propensity scores (RAMS), and inverse probability of treatment weighting (IPTW) with multinomial logistic regression or generalized boosted models.
result BART and RAMS provide lower bias and mean squared error compared to IPTW methods.
Reassesses calibration metrics in machine learning models.
problem Inconsistent reporting of calibration metrics in recent literature.
method Calibration-based decomposition of Bregman divergences, visualization of calibration and generalization error.
result New visualization technique for detecting trade-offs between calibration and generalization.
New method improves treatment effect estimation using autoencoders and causal bridge.
problem Inferring causal effects with unobserved confounders.
method Coupling autoencoder with causal bridge to estimate treatment effects.
result Improves accuracy of treatment effect estimates.
Measurement error in observational datasets can lead to systematic bias in inferences based on these datasets. As studies based on observational data are increasingly used to inform decisions with real-world impact, it is critical that we develop a robust set of techniques for analyzing and adjusting for these biases. …
Proposes a robust method for predicting missing outcomes in covariate shift adaptation.
problem Predicting missing outcomes in test data with covariate shift.
method Doubly robust estimator for covariate shift adaptation via importance weighting, incorporating an additional estimator for the regression function.
result Shows robustness against density-ratio estimation errors, maintaining consistency if either estimator is consistent.
MOB-dS uses permutation to correct for dependency in discrete survival data.
problem Identifying subgroups in discrete event time data with potential spurious results.
method Model-based recursive partitioning (MOB) with modified data matrix and permutation test.
result MOB-dS controls type I error rate better than standard MOB for discrete survival data.
This paper describes a hierarchical learning strategy for generating sparse representations of multivariate datasets. The hierarchy arises from approximation spaces considered at successively finer scales. A detailed analysis of stability, convergence and behavior of error functionals associated with the approximations…
A new meta-learner improves prediction of individualized outcomes in sequential decisions.
problem Predicting individualized outcomes over long horizons in sequential decision-making.
method Developed a novel meta-learner called DRQ-learner with theoretical guarantees of orthogonality and quasi-oracle efficiency.
result DRQ-learner achieves quasi-oracle efficiency, doubly robustness, and Neyman-orthogonality.
New calibration measure SCDL improves trust in AI predictions.
problem Improving trust in AI predictions by ensuring they are both actionable and testable.
method Introducing SCDL, a new calibration measure that is fully actionable and testable.
result SCDL is the first calibration measure that is fully actionable and testable.