RegMixMatch optimizes Mixup for semi-supervised learning by integrating high- and low-confidence samples.
problem Mixup degrades SSL performance by compromising artificial labels purity.
method RegMixMatch integrates high- and low-confidence samples, uses class-aware Mixup, and mitigates confirmation bias.
result RegMixMatch achieves state-of-the-art performance in SSL benchmarks.
A new method reduces noise in multi-label data and reduces dimensionality.
problem Handling noisy multi-label data in semi-supervised settings.
method Semi-supervised and multi-label dimensionality reduction method using label propagation.
result NMLSDR outperforms state-of-the-art algorithms in reducing noise and dimensionality.
The paper tackles fair decision-making with imperfect labels.
problem Predictive models learn from biased data due to selective labeling.
method Proposes learning decision policies that maximize utility under fairness constraints.
result Learning to decide improves fairness and utility compared to traditional risk minimization.
Proposes LatLapMED for detecting high-utility anomalies.
problem Detecting statistically rare instances with real-world significance.
method Uses EM algorithm to combine entropy minimization and maximum entropy discrimination.
result Superior performance over existing anomaly detection methods.
Graph-based multi-label classifier extends CULP for multi-label data.
problem Solving multi-label classification problems.
method Extends CULP algorithm to handle multi-label data.
result Competitive results compared to cutting-edge multi-label classifiers.
Crowdlab uses classifiers to estimate consensus labels and annotator quality.
problem Leveraging multiple annotators for data classification.
method Weighted ensemble approach using any trained classifier.
result Superior estimates for consensus labels and annotator quality.
Mitigates confirmation bias in SSL by adjusting pseudo labels dynamically.
problem Confirmation bias in semi-supervised learning leads to errors in pseudo labels.
method TaMatch framework adjusts scaling ratio to debias pseudo labels and dynamically adjusts target distribution.
result TaMatch significantly outperforms existing methods in SSL tasks.
Study active learning with imperfect labelers who can abstain or make mistakes.
problem Learning from noisy and abstaining labelers in active learning.
method Proposed an algorithm that utilizes abstention responses and analyzes its consistency and query complexity.
result Achieves nearly optimal query complexity under certain conditions.
Innovative warping labeling for twisted knots and braids.
problem Inventing an invariant for twisted knots and braids.
method Introducing warping degree, constructing warping labeling, extending to virtual braids.
result Developed invariants for twisted knots and braids using warping labeling.
Extends risk control to adaptive data collection, anytime-valid guarantees.
problem Ensuring safety of machine learning models with critical risk measures.
method Sequential risk controlling prediction sets (RCPS) for adaptive data collection and active labeling.
result Anytime-valid guarantees for risk control in sequential data collection.
Study improves deep learning chest X-ray models by incorporating lateral views.
problem Lack of lateral views in training datasets limits deep learning performance.
method Used PadChest dataset with multiple views to explore merging methods.
result Incorporating lateral views increases model performance for 32 labels.
Crowdsourcing utilizes the wisdom of crowds for collective classification via information (e.g., labels of an item) provided by labelers. Current crowdsourcing algorithms are mainly unsupervised methods that are unaware of the quality of crowdsourced data. In this paper, we propose a supervised collective classificatio…
A new multi-label classification model combining SVM and BR with low-rank learning.
problem Class imbalance and label correlation issues in multi-label classification.
method Joint Ranking SVM and Binary Relevance with robust Low-rank learning (RBRL).
result RBRL outperforms state-of-the-art methods in multi-label classification.
TA-VAAL improves active learning by better utilizing task structures and overall data distribution.
problem High labeling cost limits deep learning applications; active learning selects informative samples.
method Task-aware variational adversarial active learning (TA-VAAL) modifies VAAL by relaxing task loss prediction and using ranking loss information.
result TA-VAAL outperforms state-of-the-arts on various datasets, including balanced and imbalanced labels.
Proposes a self-paced multi-label learning method to handle diverse labels efficiently.
problem Learning from multi-label data with a large label space is NP-hard and prone to overfitting.
method Self-paced multi-label learning with diversity (SPMLD) approach, incorporating gradual label inclusion and diversity maintenance.
result The proposed SPMLD framework optimizes a non-convex objective function using block coordinate descent.
Proposes an alternative approach to propagate labels in GCNs using network diffusion and clustering.
problem Challenges of training GCNs with limited labeled data and bias in network diffusion methods.
method Clustering nodes into communities, using diffusion to quantify proximity, and comparing topological profiles.
result Identifies nodes most similar to labeled nodes, improving label propagation in GCNs.
A new method for domain generalization using unlabeled data.
problem Learning from multiple domains with limited labeled data.
method Combines meta learning and semi-supervised learning with entropy-based pseudo-labeling and discrepancy loss.
result Significantly outperforms state-of-the-art methods on benchmark datasets.
Doubly robust self-training improves semi-supervised learning by balancing labeled and pseudo-labeled data.
problem Improving semi-supervised learning performance with limited labeled data.
method Introduces doubly robust self-training, a method that combines labeled and pseudo-labeled data to balance between labeled-only and pseudo-labeled-only training.
result Demonstrates superior performance of doubly robust self-training on ImageNet and nuScenes datasets.
A new framework for federated learning tackles challenges with horizontally partitioned labels and stragglers.
problem Challenges with horizontally partitioned labels and stragglers in federated learning.
method Proposes a novel vertical federated learning framework named Cascade Vertical Federated Learning (CVFL) to fully utilize all horizontally partitioned labels and mitigate stragglers.
result Demonstrates comparable performance to centralized training and mitigates stragglers.
TCR improves DNN robustness to noisy labels with minimal overhead.
problem Training on noisy labeled datasets degrades DNN generalization.
method TCR combines original labels and previous epoch predictions for regularization.
result TCR consistently enhances DNN robustness to label noise.
New research shows label refinement and weak training have limitations for aligning LLMs.
problem Limitations of refinement methods for aligning large language models.
method Analyzed probabilistic assumptions and alternative approaches to label refinement and weak training.
result Label refinement and weak training suffer from irreducible error, leaving a performance gap.
Solves biased pseudo-labels in imbalanced SSL by refining them.
problem Imbalanced class distributions in semi-supervised learning lead to biased pseudo-labels.
method Formulates a convex optimization problem to refine pseudo-labels and develops an efficient algorithm, DARP.
result Demonstrates the effectiveness of DARP in various imbalanced semi-supervised scenarios.
Improves deep learning with less labeled data using unsupervised projection.
problem Lack of labeled data for deep learning models.
method Modified unsupervised discriminant projection as a regularization term for semi-supervised learning.
result Proposes an algorithm that enhances classification performance with minimal labeled data.
A new method combines experts' opinions to train regression models with noisy labels.
problem Training regression models with noisy labels from multiple experts.
method Estimate each labeler's expertise and combine opinions using learned weights.
result Empirically outperforms existing techniques on simulated and real data.
Method estimates dataset utility via minimal program length proxy.
problem Determining if data labels are generated by useful subroutines.
method Rissanen Data Analysis (RDA) estimates minimum description length (MDL) as a proxy.
result Method reveals dataset characteristics and utility in various NLP settings.
S2cGAN uses fewer labels to train cGANs effectively.
problem Training conditional GANs requires expensive labelled data.
method Semi-supervised training with sparse labels and unsupervised data.
result S2cGAN learns conditional mapping with sparse labels and unconditional distribution with unsupervised data.
Paper tackles super-resolving labels for weakly labeled data.
problem Real-world data scarcity with expertly labeled data.
method Nested loop with KDE to super-resolve labels.
result KDE super-resolves labels more accurately than baselines.
Simpler disentanglement method extracts correlated labels from data.
problem Disentangling factors generating data into correlated and uncorrelated parts.
method Two-step process: classifier for correlated labels, adversarial training for uncorrelated part.
result Demonstrated utility on visual and financial datasets.
The task of determining labels of all network nodes based on the knowledge about network structure and labels of some training subset of nodes is called the within-network classification. It may happen that none of the labels of the nodes is known and additionally there is no information about number of classes to whic…
DM2L tackles missing labels in multi-label learning by modeling local and global rank structures.
problem Missing labels in multi-label learning.
method DM2L imposes local low-rank structures and global high-rank structures on predictions of instances from the same and different labels, respectively.
result DM2L outperforms state-of-the-art methods in multi-label learning with missing labels.
The paper tackles optimal set prediction in multi-class classification.
problem Finding the best set of classes for uncertain predictions.
method Formalized decision-theoretic framework, quantified uncertainty, Bayes-optimal prediction algorithms.
result Efficient algorithms for optimal set prediction in multi-class classification.
Paper analyzes and improves DNNs trained with noisy labels.
problem Training deep neural networks with noisy labels.
method Characterize test accuracy as a function of noise ratio, apply cross-validation, and use Co-teaching strategy.
result Our strategy consistently improves DNNs' generalization performance.
A method uses confidence scores to handle noisy labels for each instance.
problem Learning with noisy labels where each instance's label can randomly change.
method Introduces confidence-scored instance-dependent noise (CSIDN) to estimate transition distributions for each instance.
result Demonstrates the utility and effectiveness of CSIDN through experiments with synthetic and real-world noise.
CSVAE learns interpretable latent subspaces for binary labels.
problem Learning interpretable latent representations correlated to specific labels.
method Conditional Subspace VAE (CSVAE) using mutual information minimization.
result CSVAE extracts interpretable latent subspaces for binary labels.
Methodology for visualizing labeled datasets with mixed features.
problem Visualization of labeled mixed-featured datasets.
method Developed a Max-Ratio Projection (MRP) method for continuous features and extended it to datasets with discrete and continuous features using Gaussianized distributional transforms and copula models.
result Visualization of labeled mixed-featured datasets using Max-Ratio Projection and Gaussianized distributional transforms.
Proposes methods to improve wisdom of crowds by considering worker diversity and correlations.
problem Improving wisdom of crowds by considering worker diversity and correlations.
method Proposes inference, learning, and teaching methods considering worker diversity and correlations.
result Proposes methods to improve wisdom of crowds by considering worker diversity and correlations.
TransMatch uses transfer learning to improve few-shot learning accuracy.
problem Building robust models with limited labeled data.
method Transfer-learning framework combining feature extraction, initialization, and semi-supervised learning.
result Significant improvement in few-shot learning accuracy.
SemiGNN detects financial fraud using social relations and multi-view data.
problem Detecting fraud in financial services with limited labeled data and complex interactions.
method Semi-supervised graph attentive network with hierarchical attention mechanism.
result SemiGNN achieves better accuracy on fraud detection tasks compared to state-of-the-art methods.
This paper improves multi-label classification by leveraging high-order label correlations.
problem Improving accuracy in multi-label classification tasks using label correlations.
method Exploiting high-order label correlations through a supervised learning classifier system (UCS) and label powerset (LP) strategy.
result The proposed method outperforms other LP-based methods on multiple benchmark datasets.
New framework for learning from various weak supervision types.
problem Scarcity of labeled data in real-world problems.
method Probabilistic framework based on maximum likelihood principle for deep neural networks.
result General method for learning from noisy labels, complementary labels, and coarse-grained labels.
A deep abstaining classifier tackles label noise in deep learning.
problem Label noise in deep learning training data.
method Proposes a loss function allowing deep neural networks to abstain from making predictions on confusing samples.
result Deep abstaining classifier (DAC) improves robust learning in various types of label noise.
Method converts age labels into distributions to improve speaker age estimation.
problem Label ambiguity in age labels makes precise speaker age estimation challenging.
method Converts age labels into label distributions and uses label distribution learning.
result Our method outperforms baseline methods by reducing MAE by 10% on a real-world dataset.
Detects drifts in data for classification tasks using constrained embeddings.
problem Drifts in data affect model performance; unsupervised methods ignore label information.
method Task-sensitive semi-supervised drift detection with constrained low-dimensional embedding.
result Successfully detects real drifts affecting classification performance.
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
problem Lack of gold-standard phenotype labels in EHR studies.
method Proposes Binary PheNorm, an extension that uses binary silver labels directly in phenotype scoring.
result Binary PheNorm achieved strong discrimination using binary labels alone and improved performance when combined with count labels.
Learning algorithms normally assume that there is at most one annotation or label per data point. However, in some scenarios, such as medical diagnosis and on-line collaboration,multiple annotations may be available. In either case, obtaining labels for data points can be expensive and time-consuming (in some circumsta…
Paper tackles speaker verification by removing reverberation using deep LSTM networks.
problem Improving speaker verification accuracy in reverberant environments.
method Dual-label deep LSTM networks trained to map reverberant to clean speech features.
result Evaluates performance using EERs, showing improved accuracy.
Paper estimates FPR of Bayes classifier using soft labels.
problem Determining optimal classifier performance.
method Uses soft labels and denoising technique.
result Consistent and unbiased FPR estimator developed.
DoubleMatch combines pseudo-labeling with self-supervision for SSL.
problem Lack of effective use of unlabeled data in SSL.
method Combines pseudo-labeling with self-supervised loss.
result Achieves state-of-the-art accuracies on multiple datasets.