Improved assessment of knee osteoarthritis using geodesic B-score.
problem Need for automatic, reader-independent measures of osteoarthritis clinical outcomes.
method Derive a geodesic B-score for Riemannian shape spaces, develop efficient algorithm for large shape populations.
result Geodesic B-score exhibits improved discrimination ability over Euclidean B-score.
Bias and heterogeneity in peer assessment can lead to the issue of unfair scoring in the educational field. To deal with this problem, we propose a reference ranking method for an online peer assessment system using HodgeRank. Such a scheme provides instructors with an objective scoring reference based on mathematics.
Optimizes risk assessment tools using mixed-integer programming.
problem Challenges in healthcare risk assessment due to label scarcity and asymmetric misclassification costs.
method Jointly optimizes scoring weights and category thresholds via mixed-integer programming (MIP).
result Prevents label-scarce category collapse and achieves more accurate risk categorization.
E-scores assess LLM outputs for correctness, addressing p-hacking issues.
problem Limited principled mechanisms to assess generative model correctness.
method Use e-values to complement LLM outputs with e-scores, providing flexibility in tolerance levels.
result Achieves guarantees of correctness assessment and upper bounds size distortion.
New method assesses individual training points' privacy risk without retraining.
problem Privacy vulnerability of individual training points in membership inference attacks.
method Derives a closed-form decomposition of individual black-box MIA vulnerability, extending to deep networks.
result Proposes a surrogate score operating on last-layer representations that requires only a single trained model.
The paper assesses fairness in risk score models, focusing on epistemic value.
problem Fairness of risk score models in communicating uncertainty.
method Identified key fairness desiderata, developed metrics for quantitative assessment, and applied methodology in two case studies.
result Introduced a novel calibration error metric for meaningful comparisons between groups of different sizes.
The paper introduces ESE scores for farmers to assess climate change risks.
problem Assessing climate change risks in individual farmers' credit evaluations.
method Integrating ESG variables into joint liability models and using a mean-variance utility function.
result Optimal group sizes and individual-ESE score relationships under various climatic conditions.
This guide clarifies techniques for assessing and comparing model calibration and performance.
problem Assessing and comparing the calibration and performance of predictive models in insurance and actuarial practice.
method Clarifies statistical techniques for assessing model calibration and comparing models, emphasizing the importance of specifying the prediction target functional and choosing the appropriate scoring function.
result Provides guidance for the practical choice of scoring functions and illustrates results with real data case studies.
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
New test assesses probabilistic model calibration without expensive approximations.
problem Assessing calibration of probabilistic models with scores.
method Kernel Calibration Conditional Stein Discrepancy (KCCSD) test using new score-based kernels.
result Control over type-I error with improved scalability and efficiency.
Credit scoring models support loan approval decisions in the financial services industry. Lenders train these models on data from previously granted credit applications, where the borrowers' repayment behavior has been observed. This approach creates sample bias. The scoring model (i.e., classifier) is trained on accep…
Computer-aided assessment of physical rehabilitation entails evaluation of patient performance in completing prescribed rehabilitation exercises, based on processing movement data captured with a sensory system. Despite the essential role of rehabilitation assessment toward improved patient outcomes and reduced healthc…
Framework assesses autograders' reliability and biases.
problem Mixed reliability and biases in autograders for LLM evaluation.
method Bayesian GLMs to model evaluation outcomes.
result Explicit quantification of scoring differences and biases.
MIRA scores assess conditional distribution accuracy using joint samples.
problem Assessing the accuracy of candidate conditional distributions.
method Analytic expression for Mira score based on equal probability mass regions.
result Mira enables Bayesian model comparison by quantifying alignment with true process.
Paper introduces OCRR Score for quantifying DeFi wallet credit risk.
problem Inability to assess credit risk in decentralized finance.
method Probabilistic measure based on historical and predictive on-chain activity.
result Dynamic adjustment of LTV and LT based on wallet risk profile.
Generative models assess quality on time-series data using ITS and FITD.
problem Lack of consensus for quality assessment of class-conditional generative models on time-series data.
method Introduced InceptionTime Score (ITS) and Frechet InceptionTime Distance (FITD) to evaluate generative models.
result ITS and FITD combined with TSTR can accurately assess generative model performance on time-series data.
The paper introduces a method to assess machine translation quality with confidence intervals.
problem Evaluating the uncertainty and quality of machine translation.
method Utilizes conformal predictive distributions to produce prediction intervals with guaranteed coverage.
result The method outperforms a baseline on six language pairs in terms of coverage and sharpness.
This paper argues against using calibration metrics for assessing posterior probabilities and proposes expected proper scoring rules instead.
problem The assessment of posterior probabilities generated by machine learning classifiers using calibration metrics is flawed and should be replaced with expected proper scoring rules.
method The paper reviews proper scoring rules from a practical perspective, explains why expected PSRs are a principled measure of posterior quality, and introduces a new calibration metric called calibration loss.
result Calibration loss is superior to expected calibration error and expected score divergence calibration metrics for assessing posterior probabilities.
Framework scores DeFi users based on liquidity and trading behavior.
problem Distinguishing between liquidity provision and active trading in DeFi.
method Rule-based decomposition, deep residual neural network, pool-level context.
result Deep residual neural network improves user scoring and risk assessment.
Several structure learning algorithms have been proposed towards discovering causal or Bayesian Network (BN) graphs. The validity of these algorithms tends to be evaluated by assessing the relationship between the learnt and the ground truth graph. However, there is no agreed scoring metric to determine this relationsh…
Most multi-class classifiers make their prediction for a test sample by scoring the classes and selecting the one with the highest score. Analyzing these prediction scores is useful to understand the classifier behavior and to assess its reliability. We present an interactive visualization that facilitates per-class an…
Verifying probabilistic forecasts for extreme events is a highly active research area because popular media and public opinions are naturally focused on extreme events, and biased conclusions are readily made. In this context, classical verification methods tailored for extreme events, such as thresholded and weighted …
Improves logistic regression performance with nonconvex programming.
problem Stochastic generalized linear regression with chance constraints.
method Nonconvex programming techniques, clustering, quantile estimation.
result Over 1 to 2 percent improvement in model performance.
The paper examines how machine learning tools in justice settings can unfairly affect different racial groups.
problem Machine learning tools in justice settings can unfairly affect different racial groups.
method Exploring different ideas of racial equity and their computational trade-offs.
result Computation alone is unlikely to solve the unfairness in machine learning tools for justice settings.
Automates detecting problem statements in peer assessments.
problem Identifying problem statements in peer assessment reviews.
method Used machine learning models including neural networks and traditional classifiers.
result Hierarchical Attention Network classifier achieved 93.1% accuracy.
We propose a training method for deep neural network (DNN)-based source enhancement to increase objective sound quality assessment (OSQA) scores such as the perceptual evaluation of speech quality (PESQ). In many conventional studies, DNNs have been used as a mapping function to estimate time-frequency masks and traine…
Synthetic data improves credit scoring models' performance without compromising borrower privacy.
problem Scarcity of real data for credit scoring models due to privacy concerns.
method Privacy-preserving training with synthetic data.
result Credit scoring models trained with synthetic data show a reduction of 3% in AUC and 6% in KS compared to real data models.
ScoreAG generates unrestricted adversarial images maintaining semantic integrity.
problem Limited robustness evaluations due to ℓp-norm constraints. method Score-Based Adversarial Generation (ScoreAG) using score-based generative models.
result ScoreAG improves robustness assessments across multiple benchmarks.
Improves pre-trial risk assessments by making them safer without changing existing rules.
problem Improving pre-trial risk assessments while maintaining deterministic rules.
method Developed a maximin robust optimization approach to find a safer policy.
result Can safely improve certain components of the risk assessment instrument.
FUJI scores similarity of ranked lists more robustly.
problem Improving similarity assessment of ranked lists.
method Integrates a membership function into Jaccard index for better rank consideration.
result More stable and accurate similarity estimates.
The study proposes a framework to assess sustainability of firms using fund-level classifications and portfolio holdings.
problem To capture market-based sustainability assessments of firms.
method Exploiting fund-level sustainability classifications and granular portfolio holdings to construct Market-Implied Sustainability (MIS) scores.
result MIS scores capture sustainability dimensions different from conventional ESG ratings and improve portfolio performance.
Stabilizes training of DNN for speech enhancement using PESQ scores.
problem Stability issues in training DNNs using non-differentiable OSQA scores.
method Approximate OSQA scores with a differentiable auxiliary DNN and stabilize training with reinforcement learning techniques.
result Stable training of DNN to achieve state-of-the-art PESQ scores and better sound quality.
Automated scoring engines are increasingly being used to score the free-form text responses that students give to questions. Such engines are not designed to appropriately deal with responses that a human reader would find alarming such as those that indicate an intention to self-harm or harm others, responses that all…
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models …
Protein structure prediction has been a grand challenge problem in the structure biology over the last few decades. Protein quality assessment plays a very important role in protein structure prediction. In the paper, we propose a new protein quality assessment method which can predict both local and global quality of …
A new framework assesses financial and ESG risks for sustainable investing.
problem Measuring risk and reward in sustainable investing considering environmental, social, and governance factors.
method Proposes axiomatic definitions for ESG-coherent risk measures and reward-risk ratios based on bivariate random variables.
result Empirical analysis ranks stocks using the proposed measures.
Scores measure certainty and doubt in classification predictions.
problem Quantitative uncertainty assessment in classification problems.
method Intuitive scores in Bayesian and frequentist frameworks.
result Measures assess and compare prediction quality and uncertainty.
Emergenet predicts animal influenza strain emergence, outperforming current methods.
problem Limited ability to quantitatively assess animal influenza strain emergence.
method Infer digital twin of sequence evolution using 220,151 HA sequences.
result Emergenet predictions outperform WHO seasonal vaccine recommendations and CDC IRAT scores.
SDA method reduces memory and time for assimilating noisy geophysical data.
problem Challenges in identifying state trajectories of high-dimensional geophysical systems.
method Score-based data assimilation with modified score network architecture.
result Promising results for a two-layer quasi-geostrophic model.
Paper discusses extending Gini score for tied rankings and case weights.
problem Extending Gini score for tied rankings and case weights.
method Discuss and adapt Gini score for ties and case weights.
result Gini score can be used for tied rankings and case weights.
The ROC curve is widely used to assess the quality of prediction/classification/ranking algorithms, and its properties have been extensively studied. The precision-recall (PR) curve has become the de facto replacement for the ROC curve in the presence of imbalance, namely where one class is far more likely than the oth…
Pairwise ranking aligns subjective clinical evaluations with objective indicators.
problem Aligning subjective clinical evaluations with objective indicators for improved diagnosis.
method Pairwise ranking methods to align subjective evaluations with objective indicators.
result The resulting score improves classification accuracy and provides a nuanced severity assessment.
Deconfounding scores improve causal effect estimation with weak overlap.
problem Challenges in causal treatment effect estimation due to weak overlap in high-dimensional data.
method Propose deconfounding scores to preserve identification and target estimation while improving overlap.
result Prognostic scores are overlap-optimal under a broad family of generalized linear models with Gaussian features.
New ESGM scores include a 'Missing' pillar to account for unpublished ESG data.
problem Unpublished ESG data affects the reliability of ESG scores.
method Formulated a new 'Missing' pillar and introduced ESGM scores.
result ESGM scores improve risk assessment and avoid exclusion of assets.
Identification and scoring functions are statistical tools to assess the calibration and the relative performance of risk measure estimates, e.g., in backtesting. A risk measures is called identifiable (elicitable) it it admits a strict identification function (strictly consistent scoring function). We consider measure…
GPT-4 assesses its confidence in answering USMLE questions with and without feedback.
problem Understanding AI's performance in healthcare applications, especially in sensitive areas like medical education.
method Used a prompting technique to evaluate GPT-4's confidence scores before and after answering USMLE questions, categorized into with and without feedback.
result Feedback influences relative confidence but doesn't consistently increase or decrease it.
The paper introduces a knowledge score for GPR predictions to assess their reliability.
problem Uncertainty in probabilistic predictions from Gaussian process regression models.
method A knowledge score quantifying the reduction of uncertainty in GPR predictions.
result The knowledge score improves prediction accuracy in various tasks.
A new score measures data reliability without ground truth.
problem Assessing reliability of datasets without access to ground truth.
method Define ground-truth-based orderings and propose Gram determinant score.
result Gram determinant score effectively captures data quality across diverse observation processes.