New attacks fool black-box classifiers under limited query and partial information settings.
problem Adversarial attacks on black-box neural networks with limited query access and partial information.
method Developed new attacks for query-limited, partial-information, and label-only threat models.
result Effective attacks against real-world classifiers under realistic threat models.
Highly accurate classification of network categories achieved.
problem Distinguishing between different types of networks (e.g., social vs. web graphs).
method Used a random forest classifier on both real-world and synthetic networks.
result Achieved a 94.2% classification accuracy.
Paper introduces MinDiff framework for balancing classifier performance and fairness.
problem Balancing classifier performance and fairness in machine learning models.
method MinDiff framework with kernel-based statistical dependency tests.
result Demonstrates real-world improvements in classifier performance and fairness.
In this work we present a review of the state of the art of Learning Vector Quantization (LVQ) classifiers. A taxonomy is proposed which integrates the most relevant LVQ approaches to date. The main concepts associated with modern LVQ approaches are defined. A comparison is made among eleven LVQ classifiers using one r…
Framework to mitigate adversarial attacks by allowing classifiers to abstain.
problem Vulnerability of classifiers to adversarial examples.
method Adversarial training with a rejection option.
result ATRO framework improves classifier reliability against adversarial attacks.
The paper briefly introduces multiple classifier systems and describes a new algorithm, which improves classification accuracy by means of recommendation of a proper algorithm to an object classification. This recommendation is done assuming that a classifier is likely to predict the label of the object correctly if it…
Paper tackles label noise in datasets, proposing a robust learning algorithm.
problem Real-world datasets often have label noise.
method Introduces distilled examples and proposes a learning algorithm with theoretical guarantees.
result Classifiers on distilled examples converge to the Bayes optimal classifier under certain conditions.
A new method learns from synthetic data without needing real-world examples.
problem Learning robust classifiers from limited real-world data.
method A novel setting and algorithm exploiting synthetic data independence.
result Robust classifiers trained on synthetic data generalize well to real-world domains.
diproperm tests differences in HDLSS data with binary classifiers.
problem Testing differences in HDLSS data with binary classifiers.
method DiProPerm test for binary linear classifiers.
result Validates the DiProPerm test on real-world data.
Crowdlab uses classifiers to estimate consensus labels and annotator quality.
problem Leveraging multiple annotators for data classification.
method Weighted ensemble approach using any trained classifier.
result Superior estimates for consensus labels and annotator quality.
We consider general non-Euclidean distance measures between real world objects that need to be classified. It is assumed that objects are represented by distances to other objects only. Conditions for zero-error dissimilarity based classifiers are derived. Additional conditions are given under which the zero-error deci…
New framework reduces strategic manipulation cost for minority groups in fair classification.
problem Strategic manipulation disparities in fair classification.
method Constrained optimization framework that constructs classifiers to reduce strategic manipulation cost for minority groups.
result Empirically, the approach reduces strategic manipulation cost for minority groups over multiple real-world datasets.
Suitability filter detects model performance degradation in real-world deployment.
problem Ensuring model reliability in safety-critical domains without access to ground truth labels.
method Uses suitability signals to evaluate classifier performance on unlabeled user data.
result The suitability filter reliably detects performance deviations due to covariate shift.
Meta-learning method for accurate classifier from noisy annotators' data.
problem Accurate learning from noisy labels provided by multiple annotators.
method Meta-learning neural network to embed examples in latent space and estimate annotators' abilities, then adapt classifiers using EM algorithm.
result Meta-learning method improves classifier performance with minimal labeled data.
It is common that a trained classification model is applied to the operating data that is deviated from the training data because of noise. This paper demonstrates that an ensemble classifier, Diversified Multiple Tree (DMT), is more robust in classifying noisy data than other widely used ensemble methods. DMT is teste…
Meta-learning method improves PU classification performance.
problem Improving binary classifiers from PU data in unseen tasks.
method Adapts model to PU data using related tasks and neural networks.
result Proposed method outperforms existing methods on synthetic and real-world datasets.
Bayesian model fuses multiple classifiers with explicit correlation modeling.
problem Combining outputs of multiple classifiers with explicit correlation.
method Hierarchical Bayesian model with correlated Dirichlet distribution.
result Fused classifier performance can be Bayes optimal even for highly correlated base classifiers.
We study a novel machine learning (ML) problem setting of sequentially allocating small subsets of training data amongst a large set of classifiers. The goal is to select a classifier that will give near-optimal accuracy when trained on all data, while also minimizing the cost of misallocated samples. This is motivated…
THORS converts arbitrary classifiers to cost-sensitive ones efficiently.
problem Making classifiers cost-sensitive without extensive knowledge.
method THORS uses order statistics to find optimal thresholds.
result THORS provides theoretical guarantees and lower time complexity.
A new metric estimates classifier accuracy using only training data.
problem Assessing classifier accuracy without cross-validation.
method Bayesian Area Under the ROC Curve (CBAUC) metric for linear classifiers.
result The CBAUC is faster and more accurate than conventional AUC estimators.
A new classifier updates sequentially using maximum margin principles.
problem Sequential data collection and partial labeling.
method Maximum margin classifier with Maximum Entropy Discrimination principle, kernel representation, and regularization.
result Improved performance compared to non-sequential classifiers.
Study improves quality monitoring using classifier ensembles.
problem External and internal variability in manufacturing processes makes quality control challenging.
method Proposes a proactive quality monitoring approach using classifier ensembles to predict defect occurrences.
result Ensemble classification improves accuracy compared to individual classifiers.
LoCEC classifies user relationships in large social networks, addressing sparsity issues.
problem Sparse relationship feature and label data in real social platforms.
method Local Community-based Edge Classification (LoCEC) framework with three-phase processing.
result Effective and efficient classification of user relationships in large-scale networks.
Fairness in Naive Bayes classifiers by identifying and eliminating discrimination patterns.
problem Ensuring fairness in machine learning models that use partial observations.
method Discover and eliminate discrimination patterns in naive Bayes classifiers through iterative learning.
result An algorithm that learns fair naive Bayes classifiers by removing discrimination patterns.
Estimates classifier errors without ground truth using algebraic geometry.
problem Lack of ground truth in real-world production systems.
method Non-parametric estimation using algebraic geometry to solve the self-assessment problem.
result Accuracy estimators are better than one part in a hundred.
Framework improves text classification under budget constraints.
problem Building robust text classifiers with limited computational resources.
method Jointly trains a selector to identify relevant words and passes them to a classifier, with a data aggregation scheme.
result Improves classifier performance and speeds up model with minimal accuracy loss.
Study malware evasion attacks and defenses using ML classifiers.
problem Adversarial examples can mislead ML-based malware detectors.
method Investigated white-box and grey-box evasion attacks; compared defense approaches.
result Proposed a framework for grey-box and black-box attacks.
CbMLC improves multi-label classification with noisy labels.
problem Evaluating multi-label classifiers with noisy labels.
method Context-Based Multi-Label Classifier (CbMLC) that handles noisy labels without additional supervision.
result CbMLC yields substantial improvements over previous methods in noisy label settings.
Agents learning to act autonomously in real-world domains must acquire a model of the dynamics of the domain in which they operate. Learning domain dynamics can be challenging, especially where an agent only has partial access to the world state, and/or noisy external sensors. Even in standard STRIPS domains, existing …
This study evaluates data pre-processing techniques for class imbalance in biomedical data.
problem Class imbalance in biomedical datasets affects model performance.
method Resampling and feature selection techniques evaluated using SVM, C4.5, LDA, and KNN classifiers.
result Feature Selection outperforms other methods in most cases, especially with SVM.
Retail company uses Prophet algorithm for accurate sales forecasting.
problem Accurate sales forecasting in the retail industry.
method Facebook's Prophet algorithm and backtesting strategy.
result Framework demonstrates real-world use case capabilities.
A new PU classifier PUAL tackles trifurcate data issues.
problem Training classifiers on trifurcate data containing only labeled-positive instances and unlabeled instances.
method PUAL classifier with asymmetric loss and kernel-based algorithm.
result PUAL achieves satisfactory classification on trifurcate data.
Paper proposes a method to estimate true positive proportion without knowing it.
problem Bias in binary classifier performance due to different positive item proportions.
method Maximum likelihood estimator for true proportion of positives.
result Method accurately estimates true positive proportion in data sets.
CP-GAN generates images selectively conditioned on class specificity, capturing between-class relationships.
problem Generating images selectively conditioned on class specificity in class-overlapping data.
method Proposed Classifier's Posterior GAN (CP-GAN) that redesigns generator input and objective function for class-overlapping data.
result Demonstrated effectiveness of CP-GAN using both controlled and real-world class-overlapping data.
A novel distributed adaptive NN classifier for large data sets.
problem Handling large and distributed data for efficient classification.
method Distributed adaptive nearest neighbor classifier with stochastic tuning parameter selection and early stopping rule.
result Achieves nearly optimal convergence rate under large sub-sample sizes.
The paper develops classifiers that encourage positive adaptation in machine learning settings.
problem Strategic behavior by decision subjects leads to performance loss in machine learning models.
method Formulates a two-stage game to characterize optimal strategies for model designers and decision subjects.
result Trained classifiers maintain accuracy while inducing higher improvement and less manipulation.
Study enhances classifier robustness against noisy labels.
problem Impact of label noise on model performance in real-world scenarios.
method Integrates adversarial machine learning and importance reweighting techniques with CNN.
result Improved model resilience against noisy data.
Recently, machine learning algorithms have successfully entered large-scale real-world industrial applications (e.g. search engines and email spam filters). Here, the CPU cost during test time must be budgeted and accounted for. In this paper, we address the challenge of balancing the test-time cost and the classifier …
ViewFool identifies adversarial viewpoints to test image recognition robustness.
problem Lack of robustness to viewpoint changes in visual recognition models.
method Neural Radiance Fields (NeRF) and entropic regularizer to find adversarial viewpoints.
result Common image classifiers are highly vulnerable to generated adversarial viewpoints.
Automatic anomaly detection is a major issue in various areas. Beyond mere detection, the identification of the origin of the problem that produced the anomaly is also essential. This paper introduces a general methodology that can assist human operators who aim at classifying monitoring signals. The main idea is to le…
Survey of open set recognition techniques and their limitations.
problem Recognition tasks with unknown classes during testing.
method Comprehensive review of techniques, datasets, and evaluation criteria.
result Highlighting the limitations and future directions in open set recognition.
Improves classifier accuracy in ambiguous data settings.
problem Training classifiers with partially labeled data.
method Incremental pruning of candidate labels using conformal prediction.
result Significantly improves test set accuracies of PLL classifiers.
New classifiers for HDLSS data classify without tuning, robustly.
problem Classification of high-dimensional data with small samples.
method Data-adaptive energy distance classifiers, free of tuning parameters.
result Perfect classification in HDLSS asymptotic regime under general conditions.
In many real-world applications of machine learning classifiers, it is essential to predict the probability of an example belonging to a particular class. This paper proposes a simple technique for predicting probabilities based on optimizing a ranking loss, followed by isotonic regression. This semi-parametric techniq…
PLIs improve classifier performance by fine-tuning latent representations.
problem Difficult interpretation of high-dimensional latent representations in neural networks.
method Back-propagation of manual changes to low-dimensional embeddings using t-distributed stochastic neighbourhood embeddings.
result Manual separation of class clusters in latent space enhances classifier performance.
The paper explores why adversarial attacks are inevitable for certain classifiers.
problem The inevitability of adversarial attacks on neural networks.
method Theoretical analysis and experiments on classifier robustness.
result Adversarial examples are inescapable for certain classes of problems.
Researchers create a flickering attack to fool video recognition networks.
problem Adversarial manipulation of video classification networks.
method Introducing a flickering temporal perturbation to fool video classifiers.
result Achieved high fooling ratio and temporal-invariant perturbation.
We propose an efficient method to estimate the accuracy of classifiers using only unlabeled data. We consider a setting with multiple classification problems where the target classes may be tied together through logical constraints. For example, a set of classes may be mutually exclusive, meaning that a data instance c…