The paper evaluates criteria for selecting cryptocurrencies based on historical data.
problem High risk of cryptocurrencies due to volatility.
method Characterized returns and risks using historical data in short time windows (7 and 15 days). Analyzed the importance of criteria using various methods.
result Importance of criteria for selecting cryptocurrencies is analyzed and evaluated.
Paper introduces new evaluation criteria for feature-based model explanations.
problem Establishing reliable feature importance explanations for models.
method Robustness analysis using smaller adversarial perturbations.
result New explanations that are necessary and sufficient for predictions.
In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward. Such criteria are useful for risk management, and are important in domains such as finance and process control. We propose both TD(0) and LSTD(lambda) variants with linear…
The paper proposes criteria and methods for evaluating and aggregating feature-based model explanations.
problem Lack of quantitative evaluation criteria for feature-based model explanations.
method Developed quantitative evaluation criteria (low sensitivity, high faithfulness, low complexity), devised a framework for aggregation, and derived a new aggregate Shapley value explanation function.
result A new aggregate Shapley value explanation function that minimizes sensitivity.
The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to better understand exactly what they measure, their properties and their differences. …
New method finds optimal hyperparameters for multiple tasks and criteria.
problem Finding optimal hyperparameters for multiple tasks and criteria.
method Multi-Task Multi Criteria (MTMC) method that provides Pareto-optimal solutions.
result The method selects optimal hyperparameters based on given criteria significance coefficients.
TOPNet integrates task-based evaluation into machine learning models.
problem Non-differentiable task-based evaluation criteria in real-world applications.
method Task-Oriented Prediction Network (TOPNet) with learnable surrogate loss function.
result TOPNet significantly outperforms traditional and heuristic models in financial prediction tasks.
Multi-criteria recommender systems have been increasingly valuable for helping consumers identify the most relevant items based on different dimensions of user experiences. However, previously proposed multi-criteria models did not take into account latent embeddings generated from user reviews, which capture latent se…
Paper proposes a unified sparsity-based framework for evaluating algorithmic fairness.
problem Ensuring fairness in machine learning across diverse domains.
method Unified sparsity-based framework for evaluating fairness.
result Demonstrates broad applicability and effectiveness of the framework.
The Wisdom of Crowds is a phenomenon described in social science that suggests four criteria applicable to groups of people. It is claimed that, if these criteria are satisfied, then the aggregate decisions made by a group will often be better than those of its individual members. Inspired by this concept, we present a…
When sufficient labeled data are available, classical criteria based on Receiver Operating Characteristic (ROC) or Precision-Recall (PR) curves can be used to compare the performance of un-supervised anomaly detection algorithms. However , in many situations, few or no data are labeled. This calls for alternative crite…
Much work has been done in understanding human creativity and defining measures to evaluate creativity. This is necessary mainly for the reason of having an objective and automatic way of quantifying creative artifacts. In this work, we propose a regression-based learning framework which takes into account quantitative…
A method to select validation data from a dataset using statistical criteria.
problem Selecting a validation basis from a full dataset for machine learning model validation.
method Adopting a 'design of experiments' point of view and using statistical criteria, particularly Maximum Mean Discrepancy criteria.
result The 'support points' concept is particularly relevant for selecting validation data.
Fairness in machine learning has predominantly been studied in static classification settings without concern for how decisions change the underlying population over time. Conventional wisdom suggests that fairness criteria promote the long-term well-being of those groups they aim to protect. We study how static fairne…
The paper introduces new metrics for evaluating generative models of behavior.
problem Lack of quantitative evaluation criteria for unsupervised behavior discovery.
method Proposed and investigated several metrics for generative models of behavior.
result The proposed metrics correspond with biologists' intuitions and allow for model evaluation and bias understanding.
New framework for resilient bi-criteria optimization under noisy feedback.
problem Bi-criteria combinatorial optimization with noisy function evaluations.
method Introducing (α,β,δ,extttN)-resilience and developing a black-box framework. result Achieves sublinear regret and constraint violation for bi-criteria bandit problems.
This work compares data reduction criteria for online Gaussian Processes.
problem The computational complexity of Gaussian Processes limits their applicability to small datasets and streaming scenarios.
method Unified comparison of several data reduction criteria, analyzing computational complexity and reduction behavior.
result Practical guidelines for choosing a suitable data reduction criterion for online Gaussian Processes.
Multi-label classification (MLC) is an important class of machine learning problems that come with a wide spectrum of applications, each demanding a possibly different evaluation criterion. When solving the MLC problems, we generally expect the learning algorithm to take the hidden correlation of the labels into accoun…
This work explores alternative learning criteria beyond traditional risk.
problem The trade-offs of optimizing for average performance in machine learning.
method Survey and introduction of non-traditional learning criteria.
result Emphasizes the importance of considering the loss distribution for desirable performance.
A new method for multi-criteria recommender systems using graph attention networks.
problem Lack of nuanced relationships between users and items based on specific criteria.
method MDGAT, a multi-edge bipartite graph with dual attention networks and contrastive learning.
result MDGAT achieves higher accuracy in predicting item ratings compared to baseline methods.
Visual system compares and evaluates machine learning models for clinical data predictions.
problem Challenges in comparing and evaluating different machine learning models for medical predictions.
method Developed a visual analytics system to compare and evaluate multiple models' prediction criteria and consistency.
result Demonstrated the effectiveness of the visual analytics system in assisting clinicians and researchers.
In this work we study loss functions for learning and evaluating probability distributions over large discrete domains. Unlike classification or regression where a wide variety of loss functions are used, in the distribution learning and density estimation literature, very few losses outside the dominant log loss ar…
Fairmetrics evaluates fairness in ML models for specific groups.
problem Ensuring models do not produce biased outcomes for specific groups.
method User-friendly R package for evaluating group-based fairness criteria.
result Rigorous evaluation of multiple fairness metrics.
The sBIC outperforms other model selection criteria in LDA topic modeling.
problem Selecting the optimal number of topics in Latent Dirichlet Allocation (LDA) models.
method Monte Carlo simulations comparing sBIC to other criteria.
result sBIC is superior for choosing the number of topics in LDA models.
In unsupervised classification, Hidden Markov Models (HMM) are used to account for a neighborhood structure between observations. The emission distributions are often supposed to belong to some parametric family. In this paper, a semiparametric modeling where the emission distributions are a mixture of parametric distr…
Probabilistic generative models can be used for compression, denoising, inpainting, texture synthesis, semi-supervised learning, unsupervised feature learning, and other tasks. Given this wide range of applications, it is not surprising that a lot of heterogeneity exists in the way these models are formulated, trained,…
Framework for assessing explainable AI systems.
problem Lack of consensus on explainability properties.
method Survey of literature, development of taxonomy and descriptors.
result Operationalization of the framework in Explainability Fact Sheets.
Efficiency criteria improve conformal predictors' performance.
problem Improving the performance of conformal predictors.
method Learning classifiers by minimizing observed fuzziness as a training objective function.
result Conformal predictors trained by minimizing observed fuzziness perform better than traditional ones.
Bayesian active learning improves holistic educational assessments.
problem Gap between holistic CJ and criterion-based rubrics in education.
method Extends Bayesian CJ to handle multiple LO components, using entropy-based active learning.
result Enhanced predictive rankings with uncertainty estimates and quantified assessor agreement.
Proposes a new stability measure for model fitting on similar feature data sets.
problem Model fitting on data sets with similar features is challenging.
method Tuning hyperparameters in a multi-criteria fashion with predictive accuracy and feature selection stability.
result Our approach achieves similar or better predictive performance than single-criteria and stability selection approaches.
This tutorial evaluates machine learning for healthcare applications.
problem Unrealistic expectations in healthcare AI development.
method Practical guidance on performance evaluation criteria and common mistakes.
result Better understanding and avoidance of common traps in AI healthcare.
S3VDC improves DC methods for scalability, stability, and simplicity.
problem Poor scalability, instability, and lack of simplicity in DC methods.
method Four algorithmic improvements: initial γ-training, periodic β-annealing, mini-batch GMM initialization, and inverse min-max transform. S3VDC incorporates all improvements. result S3VDC outperforms state-of-the-art methods on benchmark and industrial datasets.
Paper tackles ESG rating disagreement in sustainable investing portfolios.
problem Lack of alignment between ESG ratings from different agencies affects investment decisions.
method Proposes a nonlinear optimization model reformulated as a convex quadratic program to address ESG rating disagreement.
result The proposed model can effectively manage ESG rating disagreement and improve investment decisions.
Study evaluates various regularization methods for electricity price forecasting.
problem Improving accuracy of electricity price predictions.
method Applied ten different penalty functions to two model structures in two electricity markets.
result LQ and elastic net consistently produce more accurate forecasts than other regularization types.
The paper tests if LLMs' capabilities are executed by small subnetworks (circuits).
problem Understanding how LLMs execute their capabilities.
method Formalized criteria for circuits, developed hypothesis tests, applied to six circuits.
result Synthetic circuits align with idealized properties, while Transformer circuits vary in their alignment.
This paper examines fairness and arbitrariness in bias mitigation methods.
problem Understanding how different bias mitigation strategies affect individual predictions and whether they introduce arbitrariness.
method FRAME framework to evaluate bias mitigation through five dimensions: Impact Size, Change Direction, Decision Rates, Affected Subpopulations, and Neglected Subpopulations.
result Significant differences in the behaviors of debiasing methods were exhibited, highlighting the limitations of current fairness criteria and the inherent arbitrariness in the debiasing process.
The study evaluates different parameter selection methods for Gaussian process interpolation.
problem Choosing optimal parameters for Gaussian process interpolation.
method Empirical study using scoring rules and leave-one-out selection criteria.
result The choice of model family is often more important than the selection criterion.
In sparse regression modeling via regularization such as the lasso, it is important to select appropriate values of tuning parameters including regularization parameters. The choice of tuning parameters can be viewed as a model selection and evaluation problem. Mallows' Cp type criteria may be used as a tuning param…
Unified perspective unites Bayesian optimization and active learning for efficient goal-oriented optimization.
problem Efficiently optimize expensive engineering and scientific problems with limited data.
method Unified framework linking Bayesian infill criteria and active learning criteria.
result Unified approach formalizes Bayesian infill criteria and active learning criteria.
Proposes SNML for selecting word2vec Skip-gram dimensionality.
problem Selecting optimal dimensionality for word2vec Skip-gram models.
method Information criteria (AIC, BIC, SNML) applied to SG and SG Negative Sampling models.
result SNML outperforms AIC and BIC, selecting closer optimal dimensionality.
Active learning has shown to reduce the number of experiments needed to obtain high-confidence drug-target predictions. However, in order to actually save experiments using active learning, it is crucial to have a method to evaluate the quality of the current prediction and decide when to stop the experimentation proce…
Paper proposes metrics to evaluate AI explanations without ground truth.
problem Challenges in evaluating neural network explanations without ground truth.
method Designs four metrics to evaluate explanation results.
result New insights into neural network interpretation methods.
High-dimensional predictive models, those with more measurements than observations, require regularization to be well defined, perform well empirically, and possess theoretical guarantees. The amount of regularization, often determined by tuning parameters, is integral to achieving good performance. One can choose the …
We consider the problem of estimating the set of all inputs that leads a system to some particular behavior. The system is modeled by an expensive-to-evaluate function, such as a computer experiment, and we are interested in its excursion set, i.e. the set of points where the function takes values above or below some p…
Two new criteria help understand the advantage of deep neural networks.
problem Understanding the advantage of deepening neural networks.
method Proposed two new criteria to evaluate the expressivity of functions computable by deep neural networks.
result Increasing layers is more effective than increasing units in improving the expressivity of deep neural networks.
EvaSylv software evaluates forest management with natural risk considerations.
problem Evaluating forest management under increased natural risk due to climate change.
method User-friendly software simulates forest management scenarios, integrating natural risk using a Poisson process and Faustmann approach.
result Software optimizes forest management criteria like Faustmann value and Averaged yield value.
We present an asymptotic criterion to determine the optimal number of clusters in k-means. We consider k-means as data compression, and propose to adopt the number of clusters that minimizes the estimated description length after compression. Here we report two types of compression ratio based on two ways to quantify t…
New method speeds up model selection for complex scientific tasks.
problem Exhaustive model selection is computationally infeasible for large model spaces.
method Branch-and-bound algorithm with non-monotonic criteria.
result Guaranteed identification of optimal models with significant computational speedups.