Paper explains why small-loss criterion works for learning from noisy labels.
problem Learning from noisy labels in deep learning with limited labeled data.
method Theoretical analysis and reformulation of the small-loss criterion.
result Theoretical explanation and reformulation of the small-loss criterion.
INN method refines clean labeled data from noisy labels.
problem Handling noisy labels in deep neural networks.
method INN method based on memorization effect at neighbor regions.
result INN method resolves memorization effect shortcomings.
Binary classification improves with a small fraction of corrupted labels.
problem Binary classification with corrupted labels.
method Established corruption as a form of regularization and computed upper bounds on estimation error.
result Corruption is beneficial only up to a small fraction of the total sample, scaling with the square root of the sample size.
In many applications the process of generating label information is expensive and time consuming. We present a new method that combines active and semi-supervised deep learning to achieve high generalization performance from a deep convolutional neural network with as few known labels as possible. In a setting where a …
A scalable graph-based SSL method for large-scale data with few labels.
problem Challenges in semi-supervised learning with limited labeled data and large unlabeled data.
method Constructs a graph from a small set of high-dense vertexes to learn relationships and improve performance.
result Achieves good classification performance, especially with few labels.
Cross-prediction improves inference from small labeled datasets.
problem Valid inference from small labeled datasets with imperfect predictions.
method Imputes missing labels via machine learning and debiases predictions.
result Inferences achieve desired error probability and are more powerful.
Learning with noisy labels is one of the hottest problems in weakly-supervised learning. Based on memorization effects of deep neural networks, training on small-loss instances becomes very promising for handling noisy labels. This fosters the state-of-the-art approach "Co-teaching" that cross-trains two deep neural ne…
Unsupervised meta-learning improves learning from small labeled data.
problem Acquiring representations from unlabeled data for effective downstream learning.
method Develops an unsupervised meta-learning method that optimizes for task learning ability from unlabeled data.
result Simple task construction mechanisms, like clustering embeddings, lead to good performance on various downstream tasks.
New method pools labels from similar data items to improve learning from small samples.
problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.
New methods improve inference with scarce labels using regression.
problem Efficient inference with limited labeled data.
method Relates PPI++ to ordinary least squares regression and uses robust regressors.
result Improved variance in estimators for few-label scenarios.
Anytime-valid confirmation of label-shift corrections
problem Small-batch scientific deployments with scarce labeled outcomes
method Conditional e-value and martingale-based rule
result Nonnegative martingale and anytime-valid confirmation rule
Online active regression algorithms minimize label queries for efficient data regression.
problem Efficiently predict data points with minimal label queries in an online setting.
method Proposed online algorithms for active regression under ℓ_p loss, achieving (1+ε)-approximation with minimal label queries.
result Achieves (1+ε)-approximation with only ε^(-1) d log(nκ) label queries, matching offline methods in performance.
Method generates prototypes from small datasets for efficient learning.
problem Efficiently learning from small datasets with soft labels.
method Modular method for generating soft-label prototypical lines and Hierarchical Soft-Label Prototype k-Nearest Neighbor algorithm.
result High classification accuracy with significantly fewer prototypes than classes.
Self-training improves generalization by fitting to reliable pseudo-labels or gradually improving the classification plane.
problem Understanding how self-training improves generalization in high-dimensional Gaussian mixtures.
method Analyzing iterative self-training on binary Gaussian mixtures in the asymptotic limit.
result ST improves generalization by fitting to reliable pseudo-labels or gradually improving the classification plane.
This paper uses synthetic data to improve machine learning performance on small, imbalanced datasets.
problem Improving machine learning performance on small and imbalanced datasets.
method Generates synthetic data through convex combination and uses it in a semi-supervised learning framework with support vector machines.
result Synthetic data over-sampling supports the cluster assumption in semi-supervised learning, leading to outstanding results for small high-dimensional datasets and imbalanced learning problems.
Deep RL detects anomalies from few labeled examples and large unlabeled data.
problem Anomaly detection with limited labeled data and large unlabeled data.
method Deep reinforcement learning to optimize detection of labeled and unlabeled anomalies.
result Significantly outperforms state-of-the-art methods on 48 real-world datasets.
Bayesian SSR on graphs improves regression with noisy labels.
problem Estimating function values on graphs from noisy labeled data.
method Bayesian approach using graph Laplacian and Gaussian prior.
result Rates of contraction of posterior measure around ground truth.
Paper tackles inherent risk scoring with choice-based data labeling and synthetic data collection.
problem Inconsistent expert judgments and lack of labeled data in inherent risk scoring.
method Choice-based data labeling and synthetic data collection.
result System achieves 89% accuracy on a test set of 52 examples.
One significant challenge to scaling entity resolution algorithms to massive datasets is understanding how performance changes after moving beyond the realm of small, manually labeled reference datasets. Unlike traditional machine learning tasks, when an entity resolution algorithm performs well on small hold-out datas…
In this paper, we propose a method for training neural networks when we have a large set of data with weak labels and a small amount of data with true labels. In our proposed model, we train two neural networks: a target network, the learner and a confidence network, the meta-learner. The target network is optimized to…
SPREV simplifies visualization of complex labeled datasets.
problem Challenges of reducing dimensions and visualizing labeled datasets with small class size, high dimensionality, and low sample size.
method SPREV uses a novel dimensionality reduction technique integrating geometric principles.
result SPREV effectively visualizes hidden patterns in complex labeled datasets.
Deep semi-supervised anomaly detection improves fraud detection in financial markets.
problem Detecting fraud in high-frequency financial data with limited labeled examples.
method Evaluation of Deep Semi-Supervised Anomaly Detection (Deep SAD) on proprietary limit order book data.
result Deep SAD significantly improves fraud detection accuracy with minimal labeled data.
The paper analyzes consistency of graph-based semi-supervised learning methods for binary and multi-class classification.
problem Consistency of semi-supervised learning algorithms on graphs with noisy labels and well-clustered unlabelled data.
method The study examines graph-based probit and one-hot encoding methods for binary and multi-class classification, analyzing the consistency of optimization-based techniques.
result The analysis reveals insights into the rational function choice for optimization, improving the consistency of semi-supervised learning algorithms.
Paper tackles label noise in large datasets, purifying noisy data with a nonparametric framework.
problem Label noise in large-scale datasets with coarse labels.
method Develops a model-agnostic nonparametric framework for classification.
result Framework purifies noisy data using a small clean dataset and manages ambiguous samples.
Improves NILM with multi-label SRC, outperforming state-of-the-art.
problem Non-intrusive load monitoring (NILM) for energy disaggregation.
method Modified multi-label sparse representation based classification (SRC).
result Significant improvement over state-of-the-art techniques with minimal training data.
PSDR improves robustness against noisy labels by penalizing KL divergence between similar inputs.
problem Robust training of DNNs in datasets with noisy labels.
method Introduces PSDR, a manifold regularizer that penalizes KL divergence between similar inputs.
result Significantly improves robustness against noisy labels on benchmark datasets.
Human labeling of data can be very time-consuming and expensive, yet, in many cases it is critical for the success of the learning process. In order to minimize human labeling efforts, we propose a novel active learning solution that does not rely on existing sources of unlabeled data. It uses a small amount of labeled…
New framework improves fairness in small data settings.
problem Ensuring fairness in low-data environments.
method Combines posterior sampling exploration with fair classification.
result Framework maximizes accuracy while meeting fairness constraints.
New method selects data for labeling in RKHS to improve regression accuracy.
problem Labeling cost in supervised learning.
method Importance labeling scheme in RKHS with gradient descent.
result Gradient descent with proposed labeling scheme achieves optimal convergence rate.
PPI++ uses machine learning predictions to improve inference from small datasets.
problem Efficient inference from small labeled datasets with high-quality predictions.
method Adapts prediction-powered inference (PPI) to compute confidence sets for any parameter dimensionality.
result Improves classical intervals using only labeled data, always yielding better results.
Training deep neural networks requires massive amounts of training data, but for many tasks only limited labeled data is available. This makes weak supervision attractive, using weak or noisy signals like the output of heuristic methods or user click-through data for training. In a semi-supervised setting, we can use a…
A new method classifies high-dimensional images with minimal labels using diffusion geometry.
problem Classifying high-dimensional images efficiently and accurately.
method Spatially-regularized nonlinear diffusion geometry for clustering and active learning.
result High-accuracy labelings achieved with a very small number of training labels.
Paper proposes a framework to improve weakly supervised learning performance.
problem Weakly supervised data often lead to poor performance due to unreliable labels.
method Guides label quality optimization using a small validation set.
result Framework achieves impressive performance gains with minimal validation data.
Study presents a dataset and methods to handle noisy labels in sound event classification.
problem Label noise in sound event classification datasets.
method Developed a dataset with noisy labels and evaluated CNN baseline systems.
result Training with large amounts of noisy data can outperform training with carefully-labeled data.
Paper tackles super-resolving labels for weakly labeled data.
problem Real-world data scarcity with expertly labeled data.
method Nested loop with KDE to super-resolve labels.
result KDE super-resolves labels more accurately than baselines.
GGAN improves audio representation learning with fewer labels.
problem Learning representations for specific tasks from unlabelled data.
method Guided Generative Adversarial Neural Network (GGAN).
result GGAN learns better representations with fewer labelled data.
Paper proposes a method to adapt classifiers using complementary labels instead of true labels.
problem Training classifiers with true labels from the source domain is costly and sometimes impossible.
method Proposes a novel setting with complementary labels and a complementary label adversarial network (CLARINET).
result CLARINET significantly outperforms baselines on handwritten digits and object recognition tasks.
TrustNet robustly learns noise patterns from trusted data to improve weakly-supervised classification.
problem Robustness to label noise in weakly-supervised learning.
method TrustNet learns noise patterns from trusted data, then trains a robust classifier using these patterns.
result TrustNet outperforms state-of-the-art methods in robustness to various noise patterns.
Class imbalance is an intrinsic characteristic of multi-label data. Most of the labels in multi-label data sets are associated with a small number of training examples, much smaller compared to the size of the data set. Class imbalance poses a key challenge that plagues most multi-label learning methods. Ensemble of Cl…
Active learning method for high-dimensional data using diffusion processes.
problem High-dimensional data labeling with limited labels.
method Learning intrinsic data geometries through diffusion processes on graphs, using diffusion distances to parametrize low-dimensional structures.
result The method achieves high-accuracy labelings with only a small number of carefully chosen labels.
Paper proposes semi-supervised learning for EEG analysis.
problem Reducing workload and delays in analyzing large unlabeled EEG datasets.
method Semi-supervised deep learning algorithm using minimal labeled data.
result Predictions can be made with as little as 5 labeled examples.
Proposes a strategy to train models with minimal labeled data.
problem Scarce and expensive labeled data for medical tasks.
method Recursive training strategy to use image-level annotations for pixel-level segmentation.
result Improved segmentation of intracranial hemorrhage in CT scans.
Discriminative clustering learns from both labeled and unlabeled data.
problem Clustering complex datasets with limited labeled data.
method Gradient-based stochastic training and optimal transport with entropic regularization.
result The method can learn feature representations even in fully unsupervised settings.
Automatically mined rules from dependency parsing help neural models learn from less labeled data.
problem Lack of labeled data for aspect and opinion term extraction.
method Automatically mined rules from dependency parsing, applied to auxiliary data, combined with human-annotated data.
result Neural models achieve better performance than state-of-the-art with mined rules and auxiliary data.
This paper proposes a new method to select labeled data points using VAEs for active learning.
problem High cost of acquiring labels in supervised machine learning.
method Data-driven approach using Variational Autoencoders (VAEs) to select a diverse core-set in a low-dimensional latent space.
result Improvement in accuracy over related techniques, highlighting the representation power of generative modeling.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
A semi-supervised learning method using predefined class centroids for image classification.
problem Reducing the need for labeled data in deep learning.
method Use a small number of labeled samples and data augmentation on unlabeled samples. Constrain all samples to predefined evenly-distributed class centroids (PEDCC) using loss functions.
result Achieves state-of-the-art results with minimal labeled data.
SSNAS finds neural architectures without labeled data.
problem Limited labeled data for NAS.
method Self-supervised learning for NAS.
result Comparable results to supervised NAS with labeled data.