A new imputation method MissARF uses adversarial random forests for fast and accurate missing value imputation.
problem Handling missing values in biostatistical analyses.
method Adversarial Random Forests (ARF) for density estimation and data synthesis.
result MissARF performs comparably to state-of-the-art methods in imputation quality and runtime.
A new method for covariate shift adaptation using nearest neighbors.
problem Mitigating distribution shift between source and target datasets.
method Directly work on unlabeled target data, labeled by nearest neighbors in source data.
result Optimal choice of k=1 simplifies hyper-parameter tuning and improves efficiency. Neural network model predicts alternating event-free periods.
problem Dynamic prediction of alternating recurrent events with statistical nuance.
method Developed an online dynamic prediction framework using neural network theory.
result Outstanding performance in predicting alternating recurrent event-free time.
A new VAE model identifies and estimates treatment effects with limited overlap.
problem Identifying and estimating treatment effects when subjects with certain features belong to a single treatment group.
method Developed a latent variable model to estimate a prognostic score, which is sufficient for treatment effects. The model is a new type of VAE called β-Intact-VAE.
result The model identifies individualized treatment effects and provides TE error bounds.
Unified RL survey for healthcare AI interventions.
problem Limited real-life application of RL in healthcare.
method Unified technical survey and case studies.
result Bridge between dynamic treatment regimes and mobile health.
Bayesian Beta regression for proportions in high dimensions with theoretical guarantees.
problem Modeling bounded continuous responses in high-dimensional settings with theoretical guarantees.
method Proposes a Bayesian approach using a tempered posterior with Horseshoe prior for shrinkage and variable selection.
result Demonstrates improved estimation accuracy and model interpretability in high-dimensional scenarios.
New method for estimating and optimizing MDPs without stationarity.
problem Challenges in offline contextual MDP estimation without stationarity.
method Introduces a new adaptive estimation and cost optimization approach for contextual MDPs.
result First robust, theoretically backed method for offline contextual MDP estimation.
Study adaptive clinical trial methods for identifying patient subpopulations with treatment benefit.
problem Adaptive identification of patient subpopulations with treatment benefit in clinical trials.
method Proposes AdaGGI and AdaGCPI meta-algorithms for subpopulation construction.
result Empirical investigation of AdaGGI and AdaGCPI performance across various simulation scenarios.
ADASAP accelerates GP inference for large datasets.
problem Scaling Gaussian process inference to large datasets.
method Approximate sketch-and-project algorithm for solving linear systems.
result Sketch-and-project rapidly converges to true posterior mean.
New deep Cox mixture model improves survival analysis performance.
problem Challenges in survival analysis due to censoring and healthcare applications.
method Learning mixtures of Cox regressions with deep neural networks for hazard ratios and non-parametric baseline hazard.
result Our approach outperforms classical and modern survival analysis methods, especially in minority demographics.
New methods for selecting variables in complex biomedical data.
problem Selecting important variables in multivariate, functional, and complex biomedical data.
method Optimization-based variable selection methods for various regression models.
result Outperforms state-of-the-art methods in accuracy and speed.
The development of molecular signatures for the prediction of time-to-event outcomes is a methodologically challenging task in bioinformatics and biostatistics. Although there are numerous approaches for the derivation of marker combinations and their evaluation, the underlying methodology often suffers from the proble…
In biostatistics, propensity score is a common approach to analyze the imbalance of covariate and process confounding covariates to eliminate differences between groups. While there are an abundant amount of methods to compute propensity score, a common issue of them is the corrupted labels in the dataset. For example,…
In the recent years more and more high-dimensional data sets, where the number of parameters p is high compared to the number of observations n or even larger, are available for applied researchers. Boosting algorithms represent one of the major advances in machine learning and statistics in recent years and are su…
New method improves robustness of neural network-based debiasing.
problem Improving robustness of neural network-based debiasing.
method Moment-constrained learning for neural networks.
result Improved performance compared to state-of-the-art benchmarks.
New method uses AI predictions as cheaper alternatives to expensive outcomes.
problem Using expensive outcomes for statistical inference.
method Recalibrated prediction-powered inference using machine learning techniques.
result Significant gains in effective sample size over existing PPI proposals.
New framework estimates target functions from incomplete data.
problem Estimating target functions from partially observed data.
method IF-learning framework using influence functions.
result Two learning algorithms developed for estimation.
Teaches reproducible research to medical students and postgrads.
problem Lack of reproducibility in medical research practices.
method Designed and delivered a lecture series on reproducible research.
result Encountered practical obstacles in reproducing a published analysis.
The paper provides rigorous guarantees for m-out-of-n bootstrap estimators of sample quantiles.
problem Lack of parameter-free guarantees for robust inference with heavy-tailed data.
method Central limit theorem and Edgeworth expansion for m-out-of-n bootstrap estimators of sample quantiles.
result Established rigorous guarantees for the soundness of m-out-of-n bootstrap estimators of sample quantiles.
The paper uses graph learning to detect valid instruments in high-dimensional data for house pricing.
problem Endogeneity bias and invalid instrument validation in high-dimensional data.
method Merge variable selection algorithms and probabilistic graphs to estimate house prices and causal structure.
result Efficient data-driven instrument selection and invalid instrument purge in high-dimensional data.
Lower bound on BART's mixing time increases with data points.
problem Slow mixing time in BART's MCMC chains.
method Simplified BART with a single tree and reduced MCMC moves.
result Mixing time grows exponentially with data points.
Improved IV estimates by weighting on compliance reduces noise in treatment effect estimation.
problem Noisy IV estimates in settings with non-random treatment receipt.
method Weighting observations by estimated compliance, leveraging machine learning for compliance estimation.
result Compliance weighting reduces IV variance, improving precision of treatment effect estimates.
Tests validity of DML estimators without assumptions.
problem Validating DML estimators without making assumptions.
method Develops tests to falsify assumptions for DML estimators.
result Falsifies assumptions for DML estimators with non-trivial power.