Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4182123164 · Jun 202019922001200920182026
48 results for class-label noise

Class2Simi reduces noise in noisy label learning by transforming noisy class labels into noisy similarity labels.

problem Learning with noisy labels in supervised and unsupervised settings.
method Transforming noisy class labels into noisy similarity labels, training DNNs from noisy data pairs.
result The noise rate reduction is theoretically guaranteed, making it easier to handle noisy similarity labels.

Vote-boosting uses weighted training data to build accurate and robust ensembles.

problem Generating accurate and robust ensemble classifiers.
method Sequential ensemble learning with weighted training data and emphasis on instances with high disagreement.
result Vote-boosting is effective for generating accurate and robust ensembles, especially when noise levels are low.

KFHE uses Kalman filters to improve ensemble classification accuracy.

problem Improving multi-class ensemble classification accuracy.
method KFHE treats ensemble training as a state estimation problem using Kalman filters.
result KFHE outperforms state-of-the-art algorithms in noisy and clean datasets.

A new PLL method uses class activation values to improve robustness.

problem Weakly supervised learning with noisy data and adversarial perturbations.
method Subjective logic with class activation values for uncertainty representation and label weight re-distribution.
result More robust predictions under high noise levels, out-of-distribution examples, and adversarial perturbations.

Method reweights instances and classes to improve robustness in noisy data.

problem Improving deep learning performance in the presence of label noise.
method Formulates constrained optimization problems to assign importance weights to instances and class labels.
result Significant performance gains observed in benchmark datasets with label noise.

New framework for learning with class-conditional multi-label noise.

problem Class labels corrupted with conditional probabilities for multiple labels.
method Formalized as CCMN framework, established unbiased estimators, proved consistency with multi-label loss functions, implemented partial multi-label learning method.
result Effectiveness validated on multiple datasets and metrics.

New particle-based method improves semi-supervised learning robustness to label noise.

problem Label noise degrades semi-supervised learning accuracy.
method Particle competition and cooperation algorithm for robust semi-supervised learning.
result Improved robustness to label noise compared to existing methods.

We consider the problem of learning a measure of distance among vectors in a feature space and propose a hybrid method that simultaneously learns from similarity ratings assigned to pairs of vectors and class labels assigned to individual vectors. Our method is based on a generative model in which class labels can prov…

2012-06-29abs ↗pdf ↗

In this paper we propose a measure of clustering quality or accuracy that is appropriate in situations where it is desirable to evaluate a clustering algorithm by somehow comparing the clusters it produces with ``ground truth' consisting of classes assigned to the patterns by manual means or some other means in whose v…

2012-12-12abs ↗pdf ↗

A new method clusters covariates considering class labels for better classification.

problem Clustering covariates independently of class labels can lead to poor results.
method Formulates as convex optimization, uses ADMM for solving, and selects model via marginal likelihood.
result Proposed method offers a unique global minimum and improves classification.

Ward2ICU dataset protects patient privacy while generating synthetic ICU transitions data.

problem Protecting patient privacy while creating synthetic ICU transition data.
method Wasserstein Generative Adversarial Network (GAN) to generate synthetic data, class label balancing.
result Quality of synthetic data generation assessed through binary classification task.

A fast method for discrete OT with group-sparse regularization for class label preservation.

problem Efficiently measuring the distance between two discrete distributions with class labels.
method Fast discrete OT with group-sparse regularizers using gradient-based algorithms.
result Up to 8.6 times faster than original method without degrading accuracy.

New methods lift weak supervision to structured prediction, providing robustness guarantees.

problem Applying weak supervision techniques to structured prediction problems.
method Introducing pseudo-Euclidean embeddings, tensor decompositions, and invariants for consistent noise rate estimation.
result Generalization guarantees nearly identical to those for models trained on clean data.

Sparse coding approximates the data sample as a sparse linear combination of some basic codewords and uses the sparse codes as new presentations. In this paper, we investigate learning discriminative sparse codes by sparse coding in a semi-supervised manner, where only a few training samples are labeled. By using the m…

2013-11-26abs ↗pdf ↗

Majority Vote is optimal for reliable data labeling under certain conditions.

problem Reliable data labeling requires aggregating multiple annotators' labels, but the optimality of Majority Vote is not well understood.
method Characterized conditions under which Majority Vote achieves the optimal label estimation error.
result Majority Vote optimally recovers labels for a given class distribution under tolerable annotation noise limits.

New approach for semi-supervised learning in relational networks with different link patterns.

problem Semi-supervised learning in relational networks with heterogeneous link patterns.
method Two scalable approaches for graph-based semi-supervised learning.
result Better classification performance without prior knowledge of class interactions.

This manuscript presents some new impossibility results on adversarial robustness in machine learning, a very important yet largely open problem. We show that if conditioned on a class label the data distribution satisfies the W2W_2 Talagrand transportation-cost inequality (for example, this condition is satisfied if t…

2018-10-08abs ↗pdf ↗

VAEs struggle with surjective multimodal data, especially class labels describing images.

problem VAEs struggle to capture variability in surjective multimodal data.
method Theoretical and empirical demonstration of VAEs with a mixture of experts posterior.
result VAEs with a mixture of experts posterior can disregard variation in surjective multimodal data.

Detects object edges and assigns class labels without pixel-level annotations.

problem Semantic boundary and edge detection with image-level labels.
method Proposes a novel strategy to perform edge detection and class assignment using whole image neural nets and backpropagation.
result High pixel-wise scores indicate semantic boundary locations, suggesting edge labels are not needed during training.

Study compares BERT and XLNet for multi-class categorization of product descriptions.

problem Robustness of multi-class categorization using pre-trained contextualized language models.
method Fine-tuning BERT and XLNet on Amazon product data for multi-class classification.
result Performance decreases linearly with the number of class labels, with BERT consistently outperforming XLNet.

MPNNs struggle with class-bottlenecks and heterophily, leading to performance limitations.

problem Performance limitations of MPNNs under heterophily and structural bottlenecks.
method A statistical framework decomposing model performance into SNR components and proving bounds on sensitivity.
result Optimal graph structures for maximizing higher-order homophily are disjoint unions of single-class and two-class-bipartite clusters.

Paper tackles leveraging unlabeled data for PU classification and robust generation.

problem Scarcity of labeled data in machine learning problems.
method Introduces a novel training framework that simultaneously targets PU classification and conditional generation using extra unlabeled data.
result Proves the effectiveness of a Classifier-Noise-Invariant Conditional GAN (CNI-CGAN) that enhances PU classifier performance and leverages extra data.

Prototype selection improves DS techniques' accuracy and reduces computational cost.

problem Improving the performance of dynamic selection techniques.
method Prototype selection techniques that edit validation data to remove noise and redundant instances.
result Improves DS techniques' classification accuracy and reduces computational cost.

Paper proposes a method to train robust neural networks without labeled data.

problem Training robust neural networks without class labels.
method Adversarial contrastive learning framework using unlabeled data.
result Robust Contrastive Learning (RoCL) achieves comparable robust accuracy to supervised methods and significantly improved robustness.

Classify & Count can estimate prevalence without adjustments if optimised for quantification.

problem Estimating prevalence without adjustments using a classifier optimised for quantification.
method Classify & Count approach, optimised for quantification, local Bayes optimality.
result Optimised Classify & Count can estimate prevalence without adjustments in the binormal model.