Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

306191121 · May 202619922001200920182026
48 results for Outcome Assessment

Enhances early risk assessments for pediatric outcomes using contrastive learning.

problem Improving risk assessments in early stages of pediatric development.
method Contrastive multi-modal framework that treats each time window as a distinct modality, training on all available data.
result Consistent improvements in early-stage risk assessments validated on real-world tasks.

The study examines how to assess skill when outcomes are noisy and insufficient.

problem Determining skill when outcomes are unreliable and insufficiently numerous.
method Characterizes decision domains with noise and effective sample size, using population-level validation methods.
result Domains with noisy outcomes are unreliable for individual skill assessment.

Improved assessment of knee osteoarthritis using geodesic B-score.

problem Need for automatic, reader-independent measures of osteoarthritis clinical outcomes.
method Derive a geodesic B-score for Riemannian shape spaces, develop efficient algorithm for large shape populations.
result Geodesic B-score exhibits improved discrimination ability over Euclidean B-score.

Study examines how different assessment formats affect student learning in a data communications course.

problem Understanding how various assessment formats impact student learning outcomes.
method Comparing student learning outcomes across multiple assessment formats in a core data communications course at George Mason University.
result Collective assessment formats enhance student knowledge demonstration.

The study uses transfer learning to compare surgical outcomes across racial/ethnic subgroups.

problem Difficulty in comparing surgical outcomes due to racial/ethnic and geographic differences.
method Causal inference framework and transfer learning to incorporate data from multiple populations.
result Racial and ethnic differences in surgical outcomes are found, with non-Hispanic Black patients experiencing wide variability.

Framework assesses treatment effects by risk groups in observational studies.

problem Evaluating treatment effects in observational studies with risk stratification.
method Five-step framework for risk-based assessment of treatment effect heterogeneity.
result Low-risk patients received negligible absolute benefits, while high-risk patients had pronounced effects.

Study assesses weakly-supervised methods for rare outcomes in medical records.

problem Identifying patients with specific medical conditions using electronic health records.
method Compared three methods (PheNorm, MAP, and sureLDA) in simulations with varying outcomes and silver labels.
result No single method consistently outperformed others, but sureLDA often did well.

Develops method to assess feature importance in black-box models for unconditional distribution.

problem Lack of methods to analyze feature importance in black-box models for unconditional distribution.
method Approximation method to compute feature importance curves for unconditional distribution.
result Produces sparse and faithful results, computationally efficient.

Proposes a test to ensure predictive algorithms predict intended outcomes better than unintended ones.

problem Unintended model behavior leading to prediction of unintended outcomes.
method Falsification framework using nonparametric hypothesis testing to compare prediction losses across outcomes.
result Establishes discriminant validity with respect to gender but not race in an admissions setting.

Study validates machine learning models for patient outcomes using various methods.

problem Validating machine learning models for patient outcomes in electronic health records.
method Used three state-of-the-art machine learning methods (random forest, gradient boosting, logistic regression) to predict patient outcomes and assess feature importance.
result Permutation tests applied to random forest and gradient boosting models showed the most agreement with clinical interpretation of feature importance.

New ROC tools assess predictive abilities for any linearly ordered outcomes.

problem Fundamental restriction in ROC analysis for non-dichotomous outcomes.
method ROC movies and UROC curves for linearly ordered outcomes.
result CPA equals AUC for binary outcomes and relates to Spearman's coefficient for pairwise distinct outcomes.

This study simulates biases in classifiers to assess fairness.

problem Mitigating biases in predictive models to ensure fairness.
method Agent-based model (ABM) to generate synthetic datasets with controlled biases, applied to offline and online learning approaches.
result Demonstrates how biases in data affect classifier outcomes and how mitigations impact feature usage.

Paper uses conformal prediction sets to make criminal justice risk assessments fairer.

problem Fairness issues in criminal justice risk assessment algorithms.
method Adopting conformal prediction sets to remove unfairness from algorithms and covariates.
result Constructs confusion tables and measures fairness effectively free of racial differences.

PO-Flow models potential and counterfactual outcomes for personalized treatment decisions.

problem Predicting individualized treatment effects from observational data.
method Continuous normalizing flow (CNF) framework for causal inference.
result Unified approach to potential outcome prediction, treatment effect estimation, and counterfactual prediction.

New metric MADD assesses fairness of predictive student models.

problem Predictive student models can be biased and unfair, leading to discrimination.
method Proposes MADD metric to analyze model's discriminatory behaviors.
result Fair predictive performance does not guarantee fair behaviors or outcomes.

New method identifies proxies for causal effects on multiple outcomes.

problem Estimating causal effects in scenarios with multiple outcomes and treatments.
method Causal discovery method leveraging multiple outcomes as proxies for each treatment effect.
result Parallel studies of multiple outcomes can assist in causal identification.

The paper examines how machine learning tools in justice settings can unfairly affect different racial groups.

problem Machine learning tools in justice settings can unfairly affect different racial groups.
method Exploring different ideas of racial equity and their computational trade-offs.
result Computation alone is unlikely to solve the unfairness in machine learning tools for justice settings.

Researchers develop a model to predict MS disease progression.

problem Lack of clear understanding in predicting MS disease evolution.
method Regularized machine learning methods for binary classification and multiple output regression.
result The model accurately predicts MS disease progression from patient reported measures.

We present a novel methodology for predicting future outcomes that uses small numbers of individuals participating in an imperfect information market. By determining their risk attitudes and performing a nonlinear aggregation of their predictions, we are able to assess the probability of the future outcome of an uncert…

2001-08-02abs ↗pdf ↗

New biomarker predicts MRgFUS treatment outcome without contrast agents.

problem Inaccurate assessment of treated tissue viability after MRgFUS.
method Deep learning on noncontrast multiparametric MRI images, voxel-wise registration.
result Predicted follow-up NPV with DICE coefficient 0.71, outperforming current standard.

Improves decision making by estimating bounds on potential outcomes.

problem Estimating individual treatment effects is complex and hard to estimate.
method Developed an algorithm to learn upper and lower bounds on potential outcomes that optimize an objective function defined by the decision maker.
result Our algorithm outperforms baselines, providing tighter, more reliable bounds.

FedRD improves risk difference estimation in federated learning for clinical outcomes.

problem Privacy-preserving model co-training in medical research is hindered by server-dependent architectures and focus on relative effect measures.
method FedRD is a server-independent, communication-efficient framework for federated risk difference estimation in distributed survival data.
result FedRD provides valid confidence intervals and hypothesis testing, and is asymptotically equivalent to pooled individual-level analysis.

Models predict soccer match outcomes with similar accuracy.

problem Predicting soccer match outcomes (win, draw, loss).
method Compared Bradley-Terry extensions and hierarchical Poisson log-linear model. Parameters estimated using log-likelihood or integrated nested Laplace approximations. Predictive performance assessed using temporal validation.
result Bradley-Terry extensions and hierarchical Poisson log-linear model perform similarly in predicting match outcomes.

The study evaluates AI model performance measures for medical use.

problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.

The paper examines fairness issues in decision-making systems when protected class labels are unobserved.

problem Fairness assessment challenges when protected class labels are unavailable.
method Decomposes biases in estimating outcome disparity via threshold-based imputation and proposes a weighted estimator.
result Threshold-based imputation generally overestimates disparities, while the weighted estimator has a simpler negative bias.

Optimizes risk assessment tools using mixed-integer programming.

problem Challenges in healthcare risk assessment due to label scarcity and asymmetric misclassification costs.
method Jointly optimizes scoring weights and category thresholds via mixed-integer programming (MIP).
result Prevents label-scarce category collapse and achieves more accurate risk categorization.

The paper introduces a framework for prescriptive process monitoring that generates alarms to prevent or mitigate undesired outcomes.

problem Existing predictive process monitoring techniques do not prescribe when and how to intervene to decrease undesired outcomes.
method The paper proposes a framework that extends predictive monitoring with the ability to generate alarms, incorporating a cost model to assess the trade-off between generating alarms and the cost of undesired outcomes.
result The net cost of undesired outcomes can be minimized by optimizing the generation of alarms based on the progress of the process instance and introducing delays for triggering alarms.

The paper explores how Shapley value for a feature can vary based on model outcomes and feature distribution.

problem The uniqueness of Shapley value in explaining model predictions.
method Analyzes the relationship between feature distribution and Shapley value, and compares Shapley values for different model outcomes.
result Shapley value for a feature depends on more than just its mean and can vary significantly based on model outcome.

New methods improve off-policy evaluation for survival outcomes with censoring.

problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.

The paper proposes a method to assess surrogate heterogeneity in non-randomized data.

problem Lack of methods to evaluate surrogate heterogeneity in non-randomized data.
method Proposes a framework using meta-learners to assess surrogate heterogeneity in real-world data.
result Identifies individuals for whom the surrogate is a valid replacement of the primary outcome.

A deep learning framework assesses physical rehabilitation exercises.

problem Lack of versatile, robust, and practical assessment methods for rehabilitation exercises.
method Deep learning framework with metrics, scoring functions, and neural networks.
result First implementation of deep neural networks for rehabilitation performance assessment.

Theorem ensures superior learning outcomes for authorized learners with quantum label encoding.

problem Ensuring data security for authorized learners in machine learning.
method Quantum label encoding and PAC learning framework.
result Authorized learners achieve superior learning outcomes while eavesdroppers do not.

Develops methods for AI self-assessment to improve trustworthiness.

problem Uncertainty in AI predictions and lack of trust in AI systems.
method Uncertainty estimation techniques considering practical impacts and costs.
result Guidelines for selecting and designing effective AI self-assessment methods.

Paper compares methods for predicting early readmission in sickle-cell disease.

problem Comparing methods for early readmission prediction in a high-dimensional, heterogeneous covariates and time-to-event outcome framework.
method 8 statistical methods are compared: logistic regression, SVM, RF, GB, NN, Cox PH, CURE, C-mix models, using Elastic-Net regularization.
result C-mix model yields the best performance in both binary and survival settings.