Paper introduces new metrics for evaluating model accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper reviews metrics to assess AI model calibration accuracy.
This paper evaluates and improves metrics for identifying important features in machine learning models.
We propose a novel classifier accuracy metric: the Bayesian Area Under the Receiver Operating Characteristic Curve (CBAUC). The method estimates the area under the ROC curve and is related to the recently proposed Bayesian Error Estimator. The metric can assess the quality of a classifier using only the training datase…
LxCIM metric improves binary classification performance evaluation.
We identify and optimize the fairness-accuracy tradeoff through TAF Curves and FAUC metrics.
ID-ExpO fine-tunes neural networks for more faithful explanations.
Much of the focus in the design of deep neural networks has been on improving accuracy, leading to more powerful yet highly complex network architectures that are difficult to deploy in practical scenarios, particularly on edge devices such as mobile and other consumer devices given their high computational and memory …
New metrics needed for streaming ML due to delayed labels.
We introduce a new metric to evaluate corruption robustness of ML classifiers.
Deconfounds neural network representation similarity metrics to improve consistency and accuracy.
Metric learning makes it plausible to learn distances for complex distributions of data from labeled data. However, to date, most metric learning methods are based on a single Mahalanobis metric, which cannot handle heterogeneous data well. Those that learn multiple metrics throughout the space have demonstrated superi…
New metric space for ReLU codes connects to network safety and robustness.
New metrics CWSA and CWSA+ improve model evaluation under confidence thresholds.
Synthesizes sensor likelihoods to enforce accuracy constraints in uncertain systems.
Predicting discomfort glare in open-plan offices is a challenging problem. Although glare research has existed for more than 50 years, all current glare metrics have accuracy limitations, especially in open-plan offices with low lighting levels. Thus, it is crucial to develop a new method to predict glare more accurate…
New metric measures dynamical richness without relying on accuracy.
A number of machine learning algorithms are using a metric, or a distance, in order to compare individuals. The Euclidean distance is usually employed, but it may be more efficient to learn a parametric distance such as Mahalanobis metric. Learning such a metric is a hot topic since more than ten years now, and a numbe…
Score based learning (SBL) is a promising approach for learning Bayesian networks in the discrete domain. However, when employing SBL in the continuous domain, one is either forced to move the problem to the discrete domain or use metrics such as BIC/AIC, and these approaches are often lacking. Discretization can have …
Paper introduces negative margin loss for better few-shot classification accuracy.
In critical decision-making scenarios, optimizing accuracy can lead to a biased classifier, hence past work recommends enforcing group-based fairness metrics in addition to maximizing accuracy. However, doing so exposes the classifier to another kind of bias called infra-marginality. This refers to individual-level bia…
With the increasing availability of AI-based decision support, there is an increasing need for their certification by both AI manufacturers and notified bodies, as well as the pragmatic (real-world) validation of these systems. Therefore, there is the need for meaningful and informative ways to assess the performance o…
New method makes quality metrics scale-invariant for high-dimensional data.
In this paper, we propose a new measure to gauge the complexity of image classification problems. Given an annotated image dataset, our method computes a complexity measure called the cumulative spectral gradient (CSG) which strongly correlates with the test accuracy of convolutional neural networks (CNN). The CSG meas…
This paper investigates the theoretical foundations of metric learning, focused on three key questions that are not fully addressed in prior work: 1) we consider learning general low-dimensional (low-rank) metrics as well as sparse metrics; 2) we develop upper and lower (minimax)bounds on the generalization error; 3) w…
COMPASS improves uncertainty quantification for medical segmentation metrics.
The study improves volatility model pricing accuracy with new statistical expansions.
The paper discusses thresholds and bounds for accuracy in binary classification systems.
A commonly used evaluation metric for text-to-image synthesis is the Inception score (IS) \cite{inceptionscore}, which has been shown to be a quality metric that correlates well with human judgment. However, IS does not reveal properties of the generated images indicating the ability of a text-to-image synthesis method…
SAT improves adversarial training by smoothing the loss landscape through curriculum learning.
In machine learning, the choice of a learning algorithm that is suitable for the application domain is critical. The performance metric used to compare different algorithms must also reflect the concerns of users in the application domain under consideration. In this work, we propose a novel probability-based performan…
Unified NICEk metrics improve solar forecasting accuracy.
IVON improves LoRA finetuning with minimal overhead and significant accuracy gains.
Metric learning has been successful in learning new metrics adapted to numerical datasets. However, its development on categorical data still needs further exploration. In this paper, we propose a method, called CPML for \emph{categorical projected metric learning}, that tries to efficiently~(i.e. less computational ti…
We generalise a theorem of Engman and Abreu--Freitas on the first invariant eigenvalue of non-negatively curved -invariant metrics on to general toric Kähler metrics with non-negative scalar curvature. In particular, a simple upper bound of the first non-zero invariant eigenvalue for such metri…
Most of the work on interpretable machine learning has focused on designing either inherently interpretable models, which typically trade-off accuracy for interpretability, or post-hoc explanation systems, which lack guarantees about their explanation quality. We propose an alternative to these approaches by directly r…
Deep metric learning has been demonstrated to be highly effective in learning semantic representation and encoding information that can be used to measure data similarity, by relying on the embedding learned from metric learning. At the same time, variational autoencoder (VAE) has widely been used to approximate infere…
New scoring rules compare probabilistic top lists in classification.
This paper evaluates forecast quality in electricity markets beyond traditional accuracy measures.
EAST aligns neural network classifiers with user-defined evaluation metrics.
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
This paper visualizes uncertainty in classifier performance metrics.
New insights on Bayesian models show stochasticity doesn't always improve accuracy.
Recently, adversarial deception becomes one of the most considerable threats to deep neural networks. However, compared to extensive research in new designs of various adversarial attacks and defenses, the neural networks' intrinsic robustness property is still lack of thorough investigation. This work aims to qualitat…
Non-volatile memory, such as resistive RAM (RRAM), is an emerging energy-efficient storage, especially for low-power machine learning models on the edge. It is reported, however, that the bit error rate of RRAMs can be up to 3.3% in the ultra low-power setting, which might be crucial for many use cases. Binary neural n…
Road Network Metric Learning improves ETA prediction accuracy by addressing data sparsity.
Photo-identification (photo-id) of dolphin individuals is a commonly used technique in ecological sciences to monitor state and health of individuals, as well as to study the social structure and distribution of a population. Traditional photo-id involves a laborious manual process of matching each dolphin fin photogra…
Paper proposes a robust metric learning algorithm.