Enhances early risk assessments for pediatric outcomes using contrastive learning.
problem Improving risk assessments in early stages of pediatric development.
method Contrastive multi-modal framework that treats each time window as a distinct modality, training on all available data.
result Consistent improvements in early-stage risk assessments validated on real-world tasks.
The study examines how to assess skill when outcomes are noisy and insufficient.
problem Determining skill when outcomes are unreliable and insufficiently numerous.
method Characterizes decision domains with noise and effective sample size, using population-level validation methods.
result Domains with noisy outcomes are unreliable for individual skill assessment.
Improved assessment of knee osteoarthritis using geodesic B-score.
problem Need for automatic, reader-independent measures of osteoarthritis clinical outcomes.
method Derive a geodesic B-score for Riemannian shape spaces, develop efficient algorithm for large shape populations.
result Geodesic B-score exhibits improved discrimination ability over Euclidean B-score.
Study examines how different assessment formats affect student learning in a data communications course.
problem Understanding how various assessment formats impact student learning outcomes.
method Comparing student learning outcomes across multiple assessment formats in a core data communications course at George Mason University.
result Collective assessment formats enhance student knowledge demonstration.
The study uses transfer learning to compare surgical outcomes across racial/ethnic subgroups.
problem Difficulty in comparing surgical outcomes due to racial/ethnic and geographic differences.
method Causal inference framework and transfer learning to incorporate data from multiple populations.
result Racial and ethnic differences in surgical outcomes are found, with non-Hispanic Black patients experiencing wide variability.
Framework assesses treatment effects by risk groups in observational studies.
problem Evaluating treatment effects in observational studies with risk stratification.
method Five-step framework for risk-based assessment of treatment effect heterogeneity.
result Low-risk patients received negligible absolute benefits, while high-risk patients had pronounced effects.
New model accounts for sequential dependence in LLM reliability.
problem Uncertainty in LLM reliability assessment due to sequential interactions.
method Extended Bayesian framework with Hidden Markov Model for sequential dependence.
result Ignoring sequential dependence leads to overconfident reliability estimates.
Study assesses weakly-supervised methods for rare outcomes in medical records.
problem Identifying patients with specific medical conditions using electronic health records.
method Compared three methods (PheNorm, MAP, and sureLDA) in simulations with varying outcomes and silver labels.
result No single method consistently outperformed others, but sureLDA often did well.
Develops method to assess feature importance in black-box models for unconditional distribution.
problem Lack of methods to analyze feature importance in black-box models for unconditional distribution.
method Approximation method to compute feature importance curves for unconditional distribution.
result Produces sparse and faithful results, computationally efficient.
Proposes a test to ensure predictive algorithms predict intended outcomes better than unintended ones.
problem Unintended model behavior leading to prediction of unintended outcomes.
method Falsification framework using nonparametric hypothesis testing to compare prediction losses across outcomes.
result Establishes discriminant validity with respect to gender but not race in an admissions setting.
Study validates machine learning models for patient outcomes using various methods.
problem Validating machine learning models for patient outcomes in electronic health records.
method Used three state-of-the-art machine learning methods (random forest, gradient boosting, logistic regression) to predict patient outcomes and assess feature importance.
result Permutation tests applied to random forest and gradient boosting models showed the most agreement with clinical interpretation of feature importance.
Satellite imagery helps assess sustainable development with machine learning.
problem Lack of ground data on sustainable development outcomes.
method Combining satellite imagery with machine learning to model outcomes.
result Machine learning models perform well across multiple sustainable development domains.
New ROC tools assess predictive abilities for any linearly ordered outcomes.
problem Fundamental restriction in ROC analysis for non-dichotomous outcomes.
method ROC movies and UROC curves for linearly ordered outcomes.
result CPA equals AUC for binary outcomes and relates to Spearman's coefficient for pairwise distinct outcomes.
This study simulates biases in classifiers to assess fairness.
problem Mitigating biases in predictive models to ensure fairness.
method Agent-based model (ABM) to generate synthetic datasets with controlled biases, applied to offline and online learning approaches.
result Demonstrates how biases in data affect classifier outcomes and how mitigations impact feature usage.
Paper uses conformal prediction sets to make criminal justice risk assessments fairer.
problem Fairness issues in criminal justice risk assessment algorithms.
method Adopting conformal prediction sets to remove unfairness from algorithms and covariates.
result Constructs confusion tables and measures fairness effectively free of racial differences.
PO-Flow models potential and counterfactual outcomes for personalized treatment decisions.
problem Predicting individualized treatment effects from observational data.
method Continuous normalizing flow (CNF) framework for causal inference.
result Unified approach to potential outcome prediction, treatment effect estimation, and counterfactual prediction.
New metric MADD assesses fairness of predictive student models.
problem Predictive student models can be biased and unfair, leading to discrimination.
method Proposes MADD metric to analyze model's discriminatory behaviors.
result Fair predictive performance does not guarantee fair behaviors or outcomes.
Proposes a falsification framework to test algorithmic discriminant validity.
problem Unintended model behavior in predictive algorithms.
method Falsification framework based on statistical tests comparing prediction losses across outcomes.
result Establishes discriminant validity for some outcomes but not others.
Paper addresses fairness issues in error-prone outcomes.
problem Fairness in error-prone outcomes.
method Combining fair ML methods and measurement models.
result Using a latent variable model removes detected unfairness.
Study predicts 10-year survival rates for breast cancer patients.
problem Predicting long-term survival of breast cancer patients.
method Machine learning approaches to assess survival rates.
result Improved accuracy in predicting 10-year survival.
New method identifies proxies for causal effects on multiple outcomes.
problem Estimating causal effects in scenarios with multiple outcomes and treatments.
method Causal discovery method leveraging multiple outcomes as proxies for each treatment effect.
result Parallel studies of multiple outcomes can assist in causal identification.
This study uses OPE methods to quickly assess auction policies.
problem Rapid decision-making in dynamic auction environments.
method Off-Policy Evaluation and counterfactual methods.
result Improved policy selection and optimization.
The paper examines how machine learning tools in justice settings can unfairly affect different racial groups.
problem Machine learning tools in justice settings can unfairly affect different racial groups.
method Exploring different ideas of racial equity and their computational trade-offs.
result Computation alone is unlikely to solve the unfairness in machine learning tools for justice settings.
Researchers develop a model to predict MS disease progression.
problem Lack of clear understanding in predicting MS disease evolution.
method Regularized machine learning methods for binary classification and multiple output regression.
result The model accurately predicts MS disease progression from patient reported measures.
We present a novel methodology for predicting future outcomes that uses small numbers of individuals participating in an imperfect information market. By determining their risk attitudes and performing a nonlinear aggregation of their predictions, we are able to assess the probability of the future outcome of an uncert…
Causal ML predicts treatment outcomes, aiding personalized medicine.
problem Predicting individualized treatment effects for personalized medicine.
method Flexible, data-driven methods using causal inference with clinical trial and real-world data.
result Causal ML allows for estimating individualized treatment effects.
New biomarker predicts MRgFUS treatment outcome without contrast agents.
problem Inaccurate assessment of treated tissue viability after MRgFUS.
method Deep learning on noncontrast multiparametric MRI images, voxel-wise registration.
result Predicted follow-up NPV with DICE coefficient 0.71, outperforming current standard.
New metrics improve fairness in risk assessments.
problem Risk assessments reflect historical policies, not future decisions.
method Counterfactual analogues of metrics, doubly robust estimation.
result Fairness metrics under counterfactuals can differ from standard metrics.
Improves decision making by estimating bounds on potential outcomes.
problem Estimating individual treatment effects is complex and hard to estimate.
method Developed an algorithm to learn upper and lower bounds on potential outcomes that optimize an objective function defined by the decision maker.
result Our algorithm outperforms baselines, providing tighter, more reliable bounds.
FedRD improves risk difference estimation in federated learning for clinical outcomes.
problem Privacy-preserving model co-training in medical research is hindered by server-dependent architectures and focus on relative effect measures.
method FedRD is a server-independent, communication-efficient framework for federated risk difference estimation in distributed survival data.
result FedRD provides valid confidence intervals and hypothesis testing, and is asymptotically equivalent to pooled individual-level analysis.
Models predict soccer match outcomes with similar accuracy.
problem Predicting soccer match outcomes (win, draw, loss).
method Compared Bradley-Terry extensions and hierarchical Poisson log-linear model. Parameters estimated using log-likelihood or integrated nested Laplace approximations. Predictive performance assessed using temporal validation.
result Bradley-Terry extensions and hierarchical Poisson log-linear model perform similarly in predicting match outcomes.
The study evaluates AI model performance measures for medical use.
problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.
Cluster analysis of credit card accounts helps assess risk levels.
problem Assessing risk levels for credit accounts.
method Parametric modelling of account behavior, behavioral cluster analysis with a new dissimilarity measure.
result Interesting clusters and superior prediction of account default.
The paper examines fairness issues in decision-making systems when protected class labels are unobserved.
problem Fairness assessment challenges when protected class labels are unavailable.
method Decomposes biases in estimating outcome disparity via threshold-based imputation and proposes a weighted estimator.
result Threshold-based imputation generally overestimates disparities, while the weighted estimator has a simpler negative bias.
Optimizes risk assessment tools using mixed-integer programming.
problem Challenges in healthcare risk assessment due to label scarcity and asymmetric misclassification costs.
method Jointly optimizes scoring weights and category thresholds via mixed-integer programming (MIP).
result Prevents label-scarce category collapse and achieves more accurate risk categorization.
Patient journeys are compared to find clusters of similar disease trajectories.
problem Discovering shared health outcomes among patient journeys.
method Comparing longitudinal health data to identify clusters of similar patient trajectories.
result Clusters of patient journeys with similar health outcomes can be identified.
The paper introduces a framework for prescriptive process monitoring that generates alarms to prevent or mitigate undesired outcomes.
problem Existing predictive process monitoring techniques do not prescribe when and how to intervene to decrease undesired outcomes.
method The paper proposes a framework that extends predictive monitoring with the ability to generate alarms, incorporating a cost model to assess the trade-off between generating alarms and the cost of undesired outcomes.
result The net cost of undesired outcomes can be minimized by optimizing the generation of alarms based on the progress of the process instance and introducing delays for triggering alarms.
The paper explores how Shapley value for a feature can vary based on model outcomes and feature distribution.
problem The uniqueness of Shapley value in explaining model predictions.
method Analyzes the relationship between feature distribution and Shapley value, and compares Shapley values for different model outcomes.
result Shapley value for a feature depends on more than just its mean and can vary significantly based on model outcome.
Contrast trees assess machine learning accuracy, boosting improves performance.
problem Lack of accuracy assessment for machine learning results.
method Contrast trees and distribution boosting methods.
result Distribution boosting provides an assumption-free method for estimating outcome distributions.
New methods improve off-policy evaluation for survival outcomes with censoring.
problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.
Machine learning improves kidney transplant outcomes prediction.
problem Improving prediction of kidney transplant success.
method Random forest machine learning model trained on kidney donor risk index data.
result Random forest predicted 2,148 more successful transplants than the risk index.
The paper proposes a method to assess surrogate heterogeneity in non-randomized data.
problem Lack of methods to evaluate surrogate heterogeneity in non-randomized data.
method Proposes a framework using meta-learners to assess surrogate heterogeneity in real-world data.
result Identifies individuals for whom the surrogate is a valid replacement of the primary outcome.
A deep learning framework assesses physical rehabilitation exercises.
problem Lack of versatile, robust, and practical assessment methods for rehabilitation exercises.
method Deep learning framework with metrics, scoring functions, and neural networks.
result First implementation of deep neural networks for rehabilitation performance assessment.
Theorem ensures superior learning outcomes for authorized learners with quantum label encoding.
problem Ensuring data security for authorized learners in machine learning.
method Quantum label encoding and PAC learning framework.
result Authorized learners achieve superior learning outcomes while eavesdroppers do not.
Study finds corruption negatively impacts firm performance.
problem The impact of corruption on firm performance is examined.
method Cross-sectional data analysis of a large international dataset.
result Corruption negatively affects corporate performance.
Develops framework for valuing and assessing risk of renewable PPAs.
problem Valuation and risk assessment of non-standard renewable PPAs.
method Formalizes payoff structures, derives fair contract prices, proposes market risk-assessment methodology.
result Fair prices and risk profiles vary across technologies and contractual structures.
Develops methods for AI self-assessment to improve trustworthiness.
problem Uncertainty in AI predictions and lack of trust in AI systems.
method Uncertainty estimation techniques considering practical impacts and costs.
result Guidelines for selecting and designing effective AI self-assessment methods.
Paper compares methods for predicting early readmission in sickle-cell disease.
problem Comparing methods for early readmission prediction in a high-dimensional, heterogeneous covariates and time-to-event outcome framework.
method 8 statistical methods are compared: logistic regression, SVM, RF, GB, NN, Cox PH, CURE, C-mix models, using Elastic-Net regularization.
result C-mix model yields the best performance in both binary and survival settings.