Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
problem Average causal effect estimation with a non-differentially mismeasured binary confounder.
method Identifies conditions for proxy adjustment in the presence of a non-differentially mismeasured binary confounder.
result Adjusting for a non-differentially mismeasured binary proxy can improve estimation of the average causal effect.
Probit Monotone BART estimates binary outcomes using monotonic functions.
problem Estimating conditional mean functions for binary outcomes with monotonicity constraints.
method Proposes a new BART variant that incorporates monotonicity constraints for binary outcomes.
result Allows for more precise estimation of monotonic functions in binary outcome models.
New algorithm learns sparse GLMs for binary outcomes efficiently.
problem Sparse modeling of binary outcomes in high-dimensional data.
method Iterative hard thresholding algorithm (BIHT) for sparse GLMs.
result BIHT achieves statistical optimality for logistic regression.
A new uplift modeling approach uses binary treatment indicators more efficiently.
problem Lack of full value utilization in binary uplift modeling.
method Design a novel transformed outcome for binary target variables.
result Our new approach outperforms traditional methods in synthetic and real-world datasets.
New copula-based models for binary outcomes capture complex interactions.
problem Complex interaction effects in binary outcome models.
method Additive models with copula-based components, no discretization required.
result Better or comparable predictive performance compared to other methods.
Random forest models predict CLABSI risk in hospital admissions, with static models performing similarly to dynamic ones.
problem Predicting CLABSI risk in hospital admissions using EHR data with competing risks.
method Comparison of static and dynamic random forest models for binary, multinomial, survival, and competing risks outcomes.
result Static and dynamic random forest models perform similarly in predicting CLABSI risk, with multinomial models having the lowest computation times.
Paper proposes Bayesian TMLE methods for causal effect uncertainty quantification.
problem Quantifying uncertainty in causal effect estimation.
method Three Bayesian TMLE approaches for binary and continuous outcomes.
result BN-TMLE outperforms classical implementations in small data regimes.
This tutorial introduces causal modeling methods for researchers.
problem Understanding causal relationships in research studies.
method Integrates potential outcomes and graphical methods for causal modeling.
result Clear notation and practical examples for applied researchers.
Exact simulation of correlated binary outcomes using PMF constraints and linear programming.
problem Simulating dependent Bernoulli outcomes with specific means and correlations.
method Formulate the problem over the joint Bernoulli PMF, impose constraints, and solve as a linear program. Use convex-hull characterization and truncated-moment completion scheme for feasibility and simulation.
result Exact simulation framework for correlated binary outcomes, providing a convex-hull characterization and truncated-moment completion scheme.
Efficiently identifies important variables in binary outcomes using variational Bayes.
problem Bayesian variable selection for binary outcomes with computational challenges.
method Mean-field variational Bayes approximation with closed-form updates and efficient inference algorithm.
result Successfully identifies important variables and is orders of magnitude faster than MCMC.
Developed R package for creating nomograms for any ML algorithms.
problem Creating nomograms for any machine learning algorithms.
method Formulated a function to transform ML prediction models into nomograms, requiring specific datasets.
result Created 5 types of nomograms for various ML algorithms and predictor types.
fastHDMI improves neuroimaging variable selection in high-dimensional data.
problem Efficient variable screening in high-dimensional neuroimaging datasets.
method Three mutual information estimation methods implemented in fastHDMI.
result FFTKDE-based method superior for continuous nonlinear outcomes.
Paper introduces fair GLMs with convex penalty for equalizing GLM outcomes.
problem Achieving fairness in GLMs for practical use.
method Two fairness criteria based on GLM outcomes/log-likelihoods, achieved via a convex penalty on linear components.
result The fair GLM estimator is efficient and can handle various response variables.
Study shows how adjusting for a binary proxy can bound causal effects.
problem Bounding causal effects with a binary confounder and proxy.
method Monotonicity assumption applied to a binary confounder and observed proxy.
result Adjusting for a proxy produces a measure of the effect between unadjusted and true measures.
BELIEF framework interprets GLMs using binary linear models.
problem Understanding and interpreting generalized linear models (GLMs) with binary outcomes.
method Developed a framework called binary expansion linear effect (BELIEF) to interpret GLMs through transparent linear models.
result BELIEF framework reveals perfect predictors in complete separation scenarios.
This paper introduces a new method for uncertainty quantification in prediction models.
problem Quantifying uncertainty in high-stakes applications like medicine and finance.
method Confidence sets for outcome excursions, focusing on identifying subsets of features where outcomes exceed a threshold.
result Theoretical guarantees for the probability that confidence sets contain the true feature subset, both asymptotically and for finite sample sizes.
Background: Choosing the most performing method in terms of outcome prediction or variables selection is a recurring problem in prognosis studies, leading to many publications on methods comparison. But some aspects have received little attention. First, most comparison studies treat prediction performance and variable…
The paper explores how Shapley value for a feature can vary based on model outcomes and feature distribution.
problem The uniqueness of Shapley value in explaining model predictions.
method Analyzes the relationship between feature distribution and Shapley value, and compares Shapley values for different model outcomes.
result Shapley value for a feature depends on more than just its mean and can vary significantly based on model outcome.
Study binary choice with asymmetric loss, offering simple solutions.
problem Binary choice with asymmetric loss in data-rich environments.
method Loss-based reweighting of logistic regression or machine learning techniques.
result Valid decisions on binary outcomes with general loss functions.
Proposes MRIV framework for unbiased CATE estimation using binary IVs.
problem Bias in estimating CATEs due to unobserved confounders.
method Multiply robust machine learning framework (MRIV) for binary IVs.
result MRIV yields multiple robust convergence rates and outperforms existing methods.
Study causal inference under specific sampling methods with monotonicity assumptions.
problem Causal inference under biased sampling methods.
method Binary-outcome and binary-treatment case study with monotonicity assumptions.
result Monotonicity assumptions yield comparable results to random sampling.
gOMP algorithm selects features for various types of data.
problem Feature selection for scalable molecular data.
method Generalized Orthogonal Matching Pursuit algorithm for multiple types of data.
result gOMP performs similarly or better than LASSO on various datasets.
Throughout science and technology, receiver operating characteristic (ROC) curves and associated area under the curve (AUC) measures constitute powerful tools for assessing the predictive abilities of features, markers and tests in binary classification problems. Despite its immense popularity, ROC analysis has been su…
A new method identifies class-specific covariates in multi-class prediction tasks.
problem Identifying covariates specifically associated with one or more outcome classes in multi-class prediction tasks.
method Introducing multi forests (MuFs) with multi-way and binary splits to measure class-associated discriminatory ability.
result The multi-class VIM specifically ranks class-associated covariates highly, unlike conventional VIMs.
Kernel-based learning predicts ICU escalation from COVID-19 chest X-rays.
problem Predicting ICU escalation from chest X-rays using complex data patterns.
method Generalized Linear Models with Integrated Multiple Additive Regression with Kernels (GLIMARK).
result GLIMARK effectively predicts ICU escalation from chest X-rays.
We consider the estimation of binary election outcomes as martingales and propose an arbitrage pricing when one continuously updates estimates. We argue that the estimator needs to be priced as a binary option as the arbitrage valuation minimizes the conventionally used Brier score for tracking the accuracy of probabil…
We provide yet another proof of the existence of calibrated forecasters; it has two merits. First, it is valid for an arbitrary finite number of outcomes. Second, it is short and simple and it follows from a direct application of Blackwell's approachability theorem to carefully chosen vector-valued payoff function and …
New algorithm for nonparametric IV regression using stochastic gradients.
problem Identifying causal effects in the presence of unobservable confounders.
method Functional stochastic gradient descent for NPIV regression.
result Superior stability and competitive performance compared to existing methods.
We have developed a strategy for the analysis of newly available binary data to improve outcome predictions based on existing data (binary or non-binary). Our strategy involves two modeling approaches for the newly available data, one combining binary covariate selection via LASSO with logistic regression and one based…
Develops new methods to estimate treatment effects in survival data with competing risks.
problem Estimating treatment effects in survival data with competing risks.
method Censoring Unbiased Transformations (CUTs) for survival outcomes with and without competing risks.
result Consistent estimates of heterogeneous cumulative incidence effects and total effects using HTE learners.
New research shows fairness in machine learning can sometimes make disadvantaged groups worse off.
problem The impact of fairness constraints in machine learning on different groups.
method Unified, population-level (Bayes) framework for binary classification under prevalent group fairness notions.
result Fairness in machine learning can lead to leveling down, making one or both groups worse off.
We summarize our recent findings, where we proposed a framework for learning a Kolmogorov model, for a collection of binary random variables. More specifically, we derive conditions that link outcomes of specific random variables, and extract valuable relations from the data. We also propose an algorithm for computing …
The paper reviews and extends calibration concepts for classification and regression.
problem Formalizing compatibility between probabilistic predictions and outcomes.
method Review and extension of existing calibration concepts, introduction of new concepts.
result Hierarchical relations between calibration concepts for various data types.
Estimates causal effects using machine learning for binary treatment and mediator.
problem Estimating direct and indirect quantile treatment effects under selection-on-observables.
method Double/debiased machine learning estimators based on efficient score functions.
result Uniform consistency and asymptotic normality of effect estimators.
Federated learning approach for binary matrix factorization.
problem Efficiently factorizing binary data distributed across stakeholders while maintaining privacy.
method Proximal optimization for federated learning of relaxed binary matrix factorization.
result Federated algorithm outperforms state-of-the-art methods in quality and efficacy.
Bayesian model predicts iron deficiency from multi-source multi-way molecular data.
problem Predicting iron deficiency in rhesus monkeys from multi-source multi-way molecular data.
method Developed a Bayesian approach with a linear model incorporating multi-way dependence and varying signal sizes across sources.
result Model accurately classifies iron deficiency in monkeys and outperforms simpler models.
We present a new machine learning approach to estimate personalized treatment effects in the classical potential outcomes framework with binary outcomes. To overcome the problem that both treatment and control outcomes for the same unit are required for supervised learning, we propose surrogate loss functions that inco…
Sensitivity analysis for individualized effects in OTRs with binary risk factors.
problem Addressing omitted confounding in individualized effects of OTRs.
method Simulation-based sensitivity analysis to simulate unmeasured confounders.
result Benchmarking the strength of omitted confounding for binary risk factors.
Proposes a model for time-varying regression coefficients.
problem Uncertainty in forecasting due to changing correlations over time.
method Adopting state space literature, models how regression coefficients change over time.
result Accurate estimates for continuous outcomes but fails for binary outcomes.
We study the dynamics of co-evolution of producers and customers described by bit-strings representing individual traits. Individual ''size-like'' properties are controlled by binary encounters which outcome depends upon a recognition process. Depending upon the parameter set-up, mutual selection of producers and custo…
New split rules improve subpopulation targeting in policy-making.
problem Improving binary classification for subpopulation targeting in policy-making.
method MDFS, PFS, wEFS for maximizing distance and penalizing final splits.
result Proposed methods target more vulnerable subpopulations than classic CART/KD-CART.
Unified multitask learning framework for mixed-type outcomes.
problem Difficulty in formulating a unified objective for tasks with different outcomes.
method Multitask transformation framework with shared sparsity, using deep neural networks and rank-based optimization.
result Improved prediction and variable selection across continuous, binary, and mixed outcomes.
A method for safe online classification reduces test costs while maintaining low error rates.
problem Sequential testing for binary disease outcomes with unknown logistic model parameters.
method Joint estimation of logistic parameter and feature distribution with a conservative threshold.
result Achieves target error with high probability and requires minimal excess tests.
CausalEGM estimates causal effects by encoding confounders, improving performance in high-dimensional settings.
problem Challenges in estimating causal effects with high-dimensional confounders.
method CausalEGM framework using generative modeling to decouple confounders and estimate causal effects.
result CausalEGM outperforms existing methods in binary and continuous treatment settings, especially with large sample sizes and high-dimensional confounders.
GraphITE estimates individual effects of graph-structured treatments.
problem Estimating individual effects of complex treatment structures.
method Graph neural networks and Hilbert-Schmidt Independence Criterion regularization.
result GraphITE outperforms baselines in estimating treatment effects for large numbers of treatments.
The paper analyzes recalibration methods for binary classifiers under distribution shift.
problem Recalibrating binary classifiers to match a target prior probability.
method Analysis of distribution shift assumptions and proposal of new recalibration methods.
result QMM methods provide conservative results for risk weights functions.
Prognostic scores improve logistic regression analysis in RCTs with binary outcomes.
problem Non-collapsibility in logistic regression analysis of RCTs with binary endpoints.
method Prognostic score adjustment using AI predictions to address non-collapsibility.
result Prognostic score adjustment increases power or reduces sample size for estimating conditional odds ratios.
Proposes a new decision rule for continuous treatments.
problem Developing personalized treatment recommendations for continuous treatments.
method Jump interval-learning method to estimate conditional mean of outcomes.
result Optimal interval-valued decision rule (I2DR) for continuous treatments.