Recent advances in machine learning have led to increased deployment of black-box classifiers across a wide variety of applications. In many such situations there is a critical need to both reliably assess the performance of these pre-trained models and to perform this assessment in a label-efficient manner (given that…
Protein structure prediction has been a grand challenge problem in the structure biology over the last few decades. Protein quality assessment plays a very important role in protein structure prediction. In the paper, we propose a new protein quality assessment method which can predict both local and global quality of …
New framework assesses and benchmarks ML methods for multivariate time series.
problem Benchmarking and explaining performance of machine learning methods.
method Proposes a new framework with systematized performance-explainability characteristics.
result Illustrates application to multivariate time series classifiers.
Shai is a 10B model for asset management tasks, outperforming baselines.
problem Improving performance in asset management tasks.
method Continuous pre-training and fine-tuning on asset management-specific data.
result Shai outperforms baseline models in asset management tasks.
Enhances early risk assessments for pediatric outcomes using contrastive learning.
problem Improving risk assessments in early stages of pediatric development.
method Contrastive multi-modal framework that treats each time window as a distinct modality, training on all available data.
result Consistent improvements in early-stage risk assessments validated on real-world tasks.
Computer-aided assessment of physical rehabilitation entails evaluation of patient performance in completing prescribed rehabilitation exercises, based on processing movement data captured with a sensory system. Despite the essential role of rehabilitation assessment toward improved patient outcomes and reduced healthc…
The goal of chemmodlab is to streamline the fitting and assessment pipeline for many machine learning models in R, making it easy for researchers to compare the utility of new models. While focused on implementing methods for model fitting and assessment that have been accepted by experts in the cheminformatics field, …
Although compelling assessments have been examined in recent years, more studies are required to yield a better understanding of the several methods where assessment techniques significantly affect student learning process. Most of the educational research in this area does not consider demographics data, differing met…
To assess the performance of load disaggregation algorithms it is common practise to train a candidate algorithm on data from one or multiple households and subsequently apply cross-validation by evaluating the classification and energy estimation performance on unseen portions of the dataset derived from the same hous…
Paper proposes TDQN, a DRL strategy for optimal stock trading.
problem Optimal trading position determination in stock markets.
method Deep reinforcement learning (DRL) with Trading Deep Q-Network (TDQN) algorithm.
result TDQN strategy significantly improves Sharpe ratio performance.
Study evaluates RKHS choices for assessing graph models using KSD tests.
problem Effect of RKHS choice on KSD tests for graph model assessment.
method Investigated power performance and computational runtime of KSD tests for ERGMs and synthetic graph generators.
result Different RKHS choices affect KSD test performance and computational runtime.
Assessing generative models is not an easy task. Generative models should synthesize graphs which are not replicates of real networks but show topological features similar to real graphs. We introduce an approach for assessing graph generative models using graph classifiers. The inability of an established graph classi…
Applying deep learning methods to mammography assessment has remained a challenging topic. Dense noise with sparse expressions, mega-pixel raw data resolution, lack of diverse examples have all been factors affecting performance. The lack of pixel-level ground truths have especially limited segmentation methods in push…
The paper assesses dimensionality reduction for cryptocurrency link prediction.
problem Establishing a link between cryptocurrencies using dimensionality reduction techniques.
method Used canonical correlation analysis and principal component analysis on log returns and covariates of Bitcoin and Ethereum.
result Performance of dimensionality reduction techniques in forecasting Ethereum returns with Bitcoin features.
ABROCA assesses algorithmic bias, revealing skewed distributions that inflate results.
problem Detecting nuanced performance differences in classifier fairness.
method Study of ABROCA metric's statistical properties under various conditions.
result ABROCA distributions are skewed, inflating results by chance in imbalanced classes.
A new method assesses regression models' global optimality.
problem Challenges in evaluating regression models without access to true data.
method Information Teacher framework based on Shannon mutual information.
result Demonstrates capability to detect global optimality.
Generative models assess quality on time-series data using ITS and FITD.
problem Lack of consensus for quality assessment of class-conditional generative models on time-series data.
method Introduced InceptionTime Score (ITS) and Frechet InceptionTime Distance (FITD) to evaluate generative models.
result ITS and FITD combined with TSTR can accurately assess generative model performance on time-series data.
With the industry trend of shifting from a traditional hierarchical approach to flatter management structure, crowdsourced performance assessment gained mainstream popularity. One fundamental challenge of crowdsourced performance assessment is the risks that personal interest can introduce distortions of facts, especia…
New measure assesses deep neural networks' robustness to adversarial attacks.
problem Deep learning's fragility to adversarial attacks limits its adoption in mission-critical applications.
method Introduces residual error as a new performance measure for assessing adversarial robustness.
result Demonstrates effectiveness of residual error in assessing robustness of deep neural networks.
Study introduces new financial ratios for better predicting company performance.
problem Lack of progress in predicting company performance and assessing financial risks.
method Developed new financial and macroeconomic ratios, supervised learning models, and Bayesian models.
result New proposed variables improve model accuracy and FNN performs best across multiple tasks.
The main objective of exams consists in performing an assessment of students' expertise on a specific subject. Such expertise, also referred to as skill or knowledge level, can then be leveraged in different ways (e.g., to assign a grade to the students, to understand whether a student might need some support, etc.). S…
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
A new method evaluates invariant performance of IRM-based representations.
problem Impact of data changes on machine learning model performance.
method Proposes a novel method to evaluate invariant performance of IRM-based representations.
result Establishes a robust criterion to assess invariant performance of various representation techniques.
Study assesses additional factors for identifying persistent alpha in pension funds.
problem Identify persistent alpha in pension funds using additional factors.
method Reproduces Fama and French's (2010) experiment with additional features and compares results to 3-factor model.
result Additional factors improve persistence of alpha assessment in pension funds.
GPT-4 assesses its confidence in answering USMLE questions with and without feedback.
problem Understanding AI's performance in healthcare applications, especially in sensitive areas like medical education.
method Used a prompting technique to evaluate GPT-4's confidence scores before and after answering USMLE questions, categorized into with and without feedback.
result Feedback influences relative confidence but doesn't consistently increase or decrease it.
Compared to in-clinic balance training, in-home training is not as effective. This is, in part, due to the lack of feedback from physical therapists (PTs). Here, we analyze the feasibility of using trunk sway data and machine learning (ML) techniques to automatically evaluate balance, providing accurate assessments out…
Study assesses hyperparameter tuning for causal inference with DML.
problem Optimizing hyperparameters for causal inference with DML.
method Empirical simulation study using DML approach.
result Hyperparameter tuning crucial for causal estimation with DML.
New method approximates CV for model assessment and selection.
problem Efficient model assessment and selection with large number of folds.
method Approximates expensive refitting with a single Newton step warm-started from full training set optimizer.
result Uniform non-asymptotic, deterministic model assessment guarantees for approximate CV.
Multimodal analysis assesses job interview performance and provides feedback.
problem Assessing candidate performance in interviews for professional roles.
method Multimodal analytical framework using video, audio, and text data.
result The proposed methodology achieved promising results in predicting behavioral cues.
Flexible framework assesses multilevel data group heterogeneity.
problem Multilevel data structure complicates model selection.
method Flexible framework for assessing differences between levels of grouping variables.
result Framework reliably identifies relevant multilevel components.
PerSense assesses personality traits from text for commonsense reasoning.
problem Estimating human personality traits from text for mental health analysis.
method Aggregated Probability Density Functions (PDF) and Machine Learning (ML) models.
result PerSense algorithms achieve comparable results to ground truth data, with high accuracy for personality assessment and commonsense prediction.
5D AI model detects bad loans without biased features, improving consumer protection.
problem Detecting bad loans without biased features and improving consumer protection.
method Machine learning, BiMOPT features, European Banking Authority principles, AI principles, historical and validation datasets.
result 5D correctly detected 1,461 bad loans out of 1,613 (Sensitivity = 0.91, Prevalence = 0.0253, Positive Predictive Value = 0.19).
Study evaluates neural networks for corporate credit rating assessment.
problem Improving machine learning algorithms for credit assessment.
method Analysis of four neural network architectures (MLP, CNN, CNN2D, LSTM) on financial data from energy, financial, and healthcare sectors.
result LSTM architecture consistently outperforms others in predicting corporate credit ratings.
Study combines quantum and classical deep learning for better credit risk assessment.
problem Enhancing accuracy and efficiency in credit risk evaluation.
method Hybrid Quantum-Classical Deep Neural Network for Row-Type Dependent Predictive Analysis.
result Proposed framework enhances predictive models for different loan categories.
Fast risk assessment for autonomous vehicles using learned agent futures.
problem Risk assessment for autonomous vehicles given probabilistic predictions of other agents' futures.
method Non-sampling based methods using deep neural networks for probabilistic predictions, with Gaussian and non-Gaussian mixture models for agent positions and controls.
result Effective risk assessment for low probability events using learned models of agent futures.
Assessment of mental workload in real-world conditions is key to ensure the performance of workers executing tasks that demand sustained attention. Previous literature has employed electroencephalography (EEG) to this end despite having observed that EEG correlates of mental workload vary across subjects and physical s…
A framework assesses the quality of crowdsourced weather data.
problem Quality control and assessment of crowdsourced weather data from third-party stations.
method Proposes a simple, scalable, and interpretable AI/Stats/ML framework to assess TPAWS data.
result Demonstrates the performance of the framework using synthetic and real data.
The study evaluates AI model performance measures for medical use.
problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.
Paper proposes a natural hedging framework with graphical assessment for longevity risk management.
problem Lack of a unified framework for natural hedging and graphical risk assessment.
method Structured natural hedging framework integrated with a graphical risk metric.
result Demonstrates flexibility, interpretability, and practical value for longevity risk management.
AI systems need reliable testing to ensure safety and trustworthiness.
problem Current AI Act lacks functional trustworthiness for AI systems.
method Define technical application distribution, set risk-based performance, and conduct statistically valid testing.
result Reliable functional trustworthiness is essential for AI systems.
Paper assesses holistic risks of inference attacks on ML models.
problem Lack of comprehensive risk assessment of inference attacks on ML models.
method Presented a threat model taxonomy for four inference attacks on five model architectures and four image datasets.
result Complexity of training dataset influences attack performance; model stealing and membership inference attacks are negatively correlated.
Study identifies pitfalls in assessing hierarchies for multi-class classification.
problem Lack of understanding in selecting hierarchies for multi-class classification.
method Analyzed and compared popular approaches to extracting hierarchies.
result Hierarchy quality becomes irrelevant when using powerful classifiers.
New framework assesses deep learning models for spatio-temporal data with missing data.
problem Challenges in assessing deep learning models for spatio-temporal data with missing and heterogeneous data.
method Residual correlation analysis framework using spatio-temporal graphs and asymptotically distribution-free summary statistics.
result Identification and localization of regions where predictive performance can be improved.
Optimal allocation of human effort to correct AI assessments in decision-making.
problem How to allocate costly human effort to correct noisy or biased AI-generated assessments.
method Decision-theoretic framework treating AI assessments as signals and human judgments as costly information. Developed estimation procedures under nonparametric and linear models.
result Our approach substantially outperforms LLM-only predictions and achieves performance comparable to full human review while using only 20-30% of the human information.
In statistical classification and machine learning, as well as in social and other sciences, a number of measures of association have been proposed for assessing and comparing individual classifiers, raters, as well as their groups. In this paper, we introduce, justify, and explore several new measures of association, …
Re-speaking is a mechanism for obtaining high quality subtitles for use in live broadcast and other public events. Because it relies on humans performing the actual re-speaking, the task of estimating the quality of the results is non-trivial. Most organisations rely on humans to perform the actual quality assessment, …
The Surprise index assesses autonomous systems' competency in uncertain environments.
problem Evaluating competency of autonomous systems in dynamic, uncertain environments.
method Surprise index, a measure that quantifies system performance based on available data.
result The Surprise index can be computed for dynamic systems with Gaussian marginal distributions.
PQMass assesses generative model quality using chi-squared tests.
problem Assessing the quality of generative models without density assumptions.
method Divides sample space into regions, applies chi-squared tests to p-values.
result Effectively assesses generative model quality, novelty, and diversity.