Study assesses health plan risk measures for Solvency Capital Requirement.
problem Assessing risk measures for health plans to meet Solvency Capital Requirement.
method Three-part regression model with three GLMs for claim counts, episode allocation, and severity.
result Reduction in regression models compared to traditional methods.
Bayesian networks improve product risk assessment by handling uncertainty and causality.
problem Limited handling of uncertainty and inability to incorporate causal explanations in existing methods.
method Bayesian Networks (BNs) for improved systematic product risk assessment.
result BN approach provides more powerful and flexible risk assessments.
Unified four trade-off curves for assessing generative model proximity.
problem Quantitative assessment of proximity between two probability distributions.
method Unified four existing curves: PR, Lorenz, ROC, and Rényi divergence frontiers.
result Explicit relationship between PR and Lorenz curves with domain adaptation bounds.
Study examines how different assessment formats affect student learning in a data communications course.
problem Understanding how various assessment formats impact student learning outcomes.
method Comparing student learning outcomes across multiple assessment formats in a core data communications course at George Mason University.
result Collective assessment formats enhance student knowledge demonstration.
DeepSOFA acuity score uses deep learning for ICU patient severity assessment.
problem Static severity scoring methods are time-consuming and inaccurate for ICU patients.
method Temporal measurements and interpretable deep learning models.
result DeepSOFA yields significantly more accurate predictions of in-hospital mortality.
Paper introduces active Bayesian method for assessing black-box classifiers efficiently.
problem Need to assess performance of black-box classifiers reliably with limited labels.
method Develops inference strategies and proposes active Bayesian framework for efficient instance selection.
result Significant gains in performance assessment with fewer labels compared to traditional methods.
Improved assessment of knee osteoarthritis using geodesic B-score.
problem Need for automatic, reader-independent measures of osteoarthritis clinical outcomes.
method Derive a geodesic B-score for Riemannian shape spaces, develop efficient algorithm for large shape populations.
result Geodesic B-score exhibits improved discrimination ability over Euclidean B-score.
chemmodlab simplifies fitting and assessing machine learning models in cheminformatics.
problem Comparing the utility of new machine learning models in cheminformatics.
method Streamlines model fitting and assessment pipeline, using k-fold cross-validation and multiplicity adjustments.
result Ease of presenting statistically significant performance differences among models.
The paper assesses quality measures for machine learning models using cross-validation.
problem Evaluating the accuracy and robustness of quality measures for machine learning models.
method Cross-validation approach to estimate prediction error and quantify explained variation. Confidence bounds and local quality measures derived from residuals.
result The reliability and robustness of quality measures are assessed through numerical examples and confidence bounds.
Paper proposes a new model to assess risks in energy storage systems considering both exogenous and endogenous uncertainties.
problem Current risk assessment ignores the stochastic nature of energy storage availability.
method Data-driven unified model with exogenous and endogenous uncertainty description for four types of generic energy storage.
result Comparative results show more severe risks for endogenous uncertainty, suggesting new strategies for system operators.
The paper introduces metrics to evaluate NILM algorithms' performance on unseen buildings.
problem Assessing NILM algorithms' performance on new, unseen buildings.
method Developed several metrics to evaluate NILM algorithms' generalization ability.
result Demonstrated the utility of the proposed metrics through two case studies.
New method approximates CV for model assessment and selection.
problem Efficient model assessment and selection with large number of folds.
method Approximates expensive refitting with a single Newton step warm-started from full training set optimizer.
result Uniform non-asymptotic, deterministic model assessment guarantees for approximate CV.
Optimizes exam weights for better student assessment.
problem Designing accurate exams with generic question scores.
method Uses machine learning algorithms to adjust exam weights.
result Significant error reduction in exam scores.
Paper uses TDA for automated Parkinson's disease classification and severity assessment.
problem Manual diagnosis of neurological diseases is time-consuming and inaccurate.
method Combines Topological Data Analysis (TDA) with machine learning on postural shift data.
result Proposes a stable and accurate method for classifying Parkinson's disease.
Model uses smartphone data to assess MS trajectories.
problem Personalized longitudinal MS assessment.
method Imputation, generalized estimation equation, ensemble learning, fine-tuning.
result Promising model for predicting MS over time.
Recidivism prediction instruments (RPI's) provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the country, their use is attracting tremendous controversy. Much of the controversy c…
New correlation measures improve classifier performance assessment.
problem Improving assessment of classifiers and raters.
method Introducing CO-, ANTI-, and COANTI-correlation coefficients.
result Demonstrated new measures are powerful for classifying confusion matrices.
Develops methods for AI self-assessment to improve trustworthiness.
problem Uncertainty in AI predictions and lack of trust in AI systems.
method Uncertainty estimation techniques considering practical impacts and costs.
result Guidelines for selecting and designing effective AI self-assessment methods.
Pairwise ranking aligns subjective clinical evaluations with objective indicators.
problem Aligning subjective clinical evaluations with objective indicators for improved diagnosis.
method Pairwise ranking methods to align subjective evaluations with objective indicators.
result The resulting score improves classification accuracy and provides a nuanced severity assessment.
EBM improves car insurance claim severity and frequency prediction while maintaining interpretability.
problem Balancing predictive accuracy and interpretability in insurance claim modeling.
method Combines GAM and cyclic gradient boosting, providing interpretable predictions.
result EBM outperforms benchmark models in claim severity and frequency prediction.
Synthetic social networks closely match real-world interactions.
problem Evaluating realism of synthetic social contact networks.
method Used multiple measures of graph complexity to compare synthetic networks with stylized models and empirical data.
result Synthetic networks are more realistic than stylized models.
Study assesses forecasting models in uncertain weather data.
problem Performance of forecasting models in uncertain weather data.
method Selected influential predictors, analyzed uncertainty, trained model using observed data.
result Comparison of forecasting methods in solar PV generation.
Paper improves risk assessment and density estimation for persistence landscapes.
problem Improving risk assessment and density estimation for persistence landscapes.
method Approximate confidence intervals based on a bootstrap algorithm and kernel density estimation.
result Significant improvement in quality distribution function estimator compared to standard methods.
Study assesses linear classifiers for virus genotyping and subtyping.
problem Challenges in classifying viral sequences, especially in alignment-free methods.
method Comprehensive evaluation of linear classifiers on HCV genomes, varying parameters and sequence lengths.
result Several classifiers perform well under specific conditions, providing robust assessment.
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
New framework assesses LLM security risks in BFSI.
problem Lack of domain-specific security evaluation for LLMs in BFSI.
method Risk-aware evaluation framework combining taxonomy, automated red-teaming, and ensemble judging.
result Higher decoding stochasticity and adaptive interaction lead to more severe disclosures.
Bayesian neural networks predict AD severity from EEG data.
problem Developing low-cost, non-invasive biomarkers for AD diagnosis and progression.
method Bayesian deep neural networks using QEEG markers.
result Bayesian approach provides uncertainty bounds for AD severity prediction.
Automates detecting problem statements in peer assessments.
problem Identifying problem statements in peer assessment reviews.
method Used machine learning models including neural networks and traditional classifiers.
result Hierarchical Attention Network classifier achieved 93.1% accuracy.
The study introduces a holdout-based framework to assess synthetic data fidelity and privacy.
problem Evaluating the quality and privacy of synthetic data solutions for mixed-type tabular data.
method Holdout-based empirical assessment framework measuring fidelity and privacy risk.
result Synthetic data samples are as close to the training as to the holdout data, indicating generalization and independence from individual records.
Study develops a new tool for assessing asphalt pavement conditions using deep learning.
problem Challenges in automated pavement distress detection via road images.
method Developed a hybrid model using YOLO for classification and U-net for segmentation, creating a comprehensive pavement condition tool.
result Created a new asphalt pavement condition index using deep learning.
Improves pre-trial risk assessments by making them safer without changing existing rules.
problem Improving pre-trial risk assessments while maintaining deterministic rules.
method Developed a maximin robust optimization approach to find a safer policy.
result Can safely improve certain components of the risk assessment instrument.
The paper introduces a new risk assessment framework using φ-divergence.
problem Assessing risk and decision-making in uncertain conditions.
method Introduces a novel framework called the φ-Divergence Quadrangle.
result Provides a more nuanced understanding of risk through φ-divergence.
Machine learning assesses group collaboration in classrooms.
problem Assess student group collaboration in large classrooms.
method Deep learning models using Mixup data augmentation and ordinal-cross-entropy loss function.
result Improved assessment of group collaboration quality.
Neural nets detect alarming student responses for quick review.
problem Identifying alarming student responses in online assessments.
method Developed neural network models to flag potentially concerning responses.
result Neural nets can flag alarming responses more efficiently than manual review.
This paper assesses Gaussian and Exponential mechanisms for certifying adversarial robustness.
problem Certifying adversarial robustness using randomized smoothing mechanisms.
method Proposes a generic framework to assess the appropriateness of randomized smoothing mechanisms.
result Gaussian mechanism is an appropriate option for certifying both ℓ2-norm and ℓ∞-norm robustness. New methods for functional data analysis improve manifold methods for continuous data.
problem Challenges in evaluating embeddings for functional data.
method Transfer manifold methods from tabular and image data to functional data, define a theoretical framework, and propose nuanced evaluation strategies.
result Manifold methods can be successfully applied to functional data, but careful evaluation is needed.
Re-speaking is a mechanism for obtaining high quality subtitles for use in live broadcast and other public events. Because it relies on humans performing the actual re-speaking, the task of estimating the quality of the results is non-trivial. Most organisations rely on humans to perform the actual quality assessment, …
Testing symmetry of a probability distribution is a common question arising from applications in several fields. Particularly, in the study of observables used in the analysis of stock market index variations, the question of symmetry has not been fully investigated by means of statistical procedures. In this work a di…
Proposes CCE to assess point-wise reliability of neural network predictions.
problem Overconfidence and misaligned predictive distributions in neural networks.
method Introduces Conditional Congruence (CCE) metric using conditional kernel mean embeddings.
result CCE exhibits correctness, monotonicity, reliability, and robustness in high-dimensional regression tasks.
Study evaluates neural networks for corporate credit rating assessment.
problem Improving machine learning algorithms for credit assessment.
method Analysis of four neural network architectures (MLP, CNN, CNN2D, LSTM) on financial data from energy, financial, and healthcare sectors.
result LSTM architecture consistently outperforms others in predicting corporate credit ratings.
Researchers develop tests to assess quality of GAN-generated images.
problem Lack of objective means to evaluate domain-relevant quality of GAN-generated images.
method Designed stochastic context models (SCMs) and statistical classifiers to detect high-order spatial arrangements in GAN-generated images.
result GANs can generate images that appear accurate visually but lack specific high-order spatial arrangements.
The paper finds a pervasive and severe bias in accounting semi-identity models.
problem Bias in investment-cash flow sensitivity models.
method Augmented specification with a bias-capturing variable tested across multiple databases.
result The Accounting Semi-Identity (ASI) distortion is universal and severe, affecting 100% of databases and explaining more than 83% of total explained variance.
New measure assesses deep neural networks' robustness to adversarial attacks.
problem Deep learning's fragility to adversarial attacks limits its adoption in mission-critical applications.
method Introduces residual error as a new performance measure for assessing adversarial robustness.
result Demonstrates effectiveness of residual error in assessing robustness of deep neural networks.
Study assesses additional factors for identifying persistent alpha in pension funds.
problem Identify persistent alpha in pension funds using additional factors.
method Reproduces Fama and French's (2010) experiment with additional features and compares results to 3-factor model.
result Additional factors improve persistence of alpha assessment in pension funds.
We quantify predictive uncertainty using the posterior predictive variance.
problem Quantifying uncertainty in predictive models.
method Using the law of total variance, we generate expansions for the posterior predictive variance.
result Identify the main contributors to prediction intervals and quantify term-wise uncertainty.
Paper assesses holistic risks of inference attacks on ML models.
problem Lack of comprehensive risk assessment of inference attacks on ML models.
method Presented a threat model taxonomy for four inference attacks on five model architectures and four image datasets.
result Complexity of training dataset influences attack performance; model stealing and membership inference attacks are negatively correlated.
Deep learning improves macroeconomic forecasting and risk assessment.
problem Improving accuracy in macroeconomic forecasting and sovereign risk assessment.
method Nowcasting and forecasting using deep learning techniques.
result Deep learning methods outperform traditional econometric techniques in out-of-sample performance.
The implementation of the Own Risk and Solvency Assessment is a critical issue raised by Pillar II of Solvency II framework. In particular the Overall Solvency Needs calculation left the Insurance companies to define an optimal entity-specific solvency constraint on a multi-year time horizon. In a life insurance societ…