The paper explores how graph-based semi-supervised learning algorithms behave in large data and noiseless conditions.
problem Understanding semi-supervised learning algorithms in the large graph limit and zero noise scenario.
method The study uses graph Laplacian scaling and various optimization and Bayesian approaches to find continuum limits of these algorithms.
result Conditions are identified for well-defined continuum limits of semi-supervised learning problems.
Survey explores theoretical limits and gains of semi-supervised learning.
problem Improving supervised methods using unlabeled data.
method Analysis of different semi-supervised learning methods and their assumptions.
result Understanding the limits and potential of semi-supervised learning.
Improved self-supervised learning for document images.
problem Performance of self-supervised pre-training on document images is poor.
method Proposed context-aware alternatives and a novel multi-modal method.
result Novel method outperforms other self-supervised methods on document image classification.
Improved audio classification with limited labels using multitask and self-supervised learning.
problem Limited labeled data for audio classification.
method Multitask learning and self-supervised learning on unlabeled data.
result Significant improvement in performance (up to 6%) through multitask and self-supervised learning.
This work improves medical image segmentation with limited annotations using contrastive learning.
problem Lack of labeled data for medical image segmentation.
method Contrastive learning framework for semi-supervised segmentation with domain-specific and problem-specific cues.
result Significant improvements in segmentation performance compared to other methods.
Self-supervised method improves biosignal models with limited labeled data and subjects.
problem Limited labeled data and subjects in biosignals datasets.
method Contrastive learning with subject-aware loss and data augmentation.
result Self-supervised embeddings yield competitive results compared to supervised methods.
WSGN detects actions from weak supervision, improving performance on THUMOS14 and Charades.
problem Challenging action detection requires detailed manual supervision.
method WSGN learns action detection from video-level labels, exploiting both video-specific and dataset-wide statistics.
result WSGN achieves significant gains in action detection for THUMOS14 and Charades datasets.
Proposes a new contrastive loss for semi-supervised medical image segmentation.
problem Lack of labeled data for medical image segmentation.
method Uses pseudo-labels and a local contrastive loss to learn good local representations.
result Achieved high segmentation performance on public cardiac and prostate datasets.
Paper resolves the debate on process vs. outcome supervision in reinforcement learning.
problem Distinguishing between process and outcome supervision in reinforcement learning.
method Developed a technical tool (Change of Trajectory Measure Lemma) to show equivalence between outcome and process supervision under standard data coverage assumptions.
result Reinforcement learning through outcome supervision is statistically equivalent to process supervision, up to polynomial factors in horizon.
RKD improves clustering in semi-supervised learning with limited labels.
problem Improving clustering accuracy in semi-supervised learning with few labeled examples.
method RKD as spectral clustering on a teacher model's graph, with clustering error quantification.
result RKD provably leads to low clustering error in semi-supervised classification problems.
Improved self-supervised learning on ImageNet achieves top-1 accuracy of 77.1%.
problem Self-supervised ResNets underperform supervised learning on ImageNet.
method ReLICv2 combines explicit invariance loss with contrastive objective over varied data views.
result ReLICv2 achieves 77.1% top-1 accuracy on ImageNet, improving over previous state-of-the-art by 1.5%.
Despite significant progress in object categorization, in recent years, a number of important challenges remain, mainly, ability to learn from limited labeled data and ability to recognize object classes within large, potentially open, set of labels. Zero-shot learning is one way of addressing these challenges, but it …
Proposes a method to train deep neural networks with limited labeled data.
problem Lack of labeled data in neural text classification.
method Two modules: pseudo-document generator and self-training module.
result Significantly outperforms baseline methods without excessive labeled data.
Self-training improves GANs for semi-supervised learning.
problem Training GANs with limited labeled data.
method Combining self-training with GANs' infinite data generation.
result Self-training improves GANs' performance in semi-supervised learning.
Self-supervised method improves CBIR of CT liver images.
problem Limited labeled data and lack of transparency in deep CBIR systems.
method Proposes a self-supervised learning framework with domain-knowledge integration.
result Improved performance and generalization across datasets.
Limited supervision can enable reliable disentangled representation learning.
problem Learning disentangled representations without inductive biases is theoretically impossible.
method Investigated the impact of limited supervision (0.01--0.5% of data) on disentanglement methods.
result A small number of labeled examples (0.01--0.5\% of the data set) is sufficient for model selection.
The study explores the strengths and weaknesses of models that generalize from weak to strong supervision.
problem Understanding the limitations and capabilities of models that generalize from weak to strong supervision.
method Theoretical analysis and experimental validation in both classification and regression settings.
result Theoretical bounds reveal the importance of strong generalization and calibration of the weak model and a careful balance in the training process.
This research improves representation learning for new domains with limited new supervision.
problem Learning representations that generalize well to new domains with minimal new data.
method Encourages linearity of factors of variation through learned linear transformations called latent canonicalizers.
result Reduces the number of observations needed to generalize to a similar target domain compared to supervised baselines.
Improved VAE learns disentangled representations with less supervision.
problem Learning disentangled representations is challenging.
method Semi-supervised disentanglement learning with label replacement.
result Significant improvement in disentanglement with minimal supervision.
Paper develops methods for semi-supervised Fréchet regression.
problem High costs of obtaining non-Euclidean labels.
method Proposes semi-supervised NW Fréchet regression and semi-supervised kNN Fréchet regression.
result Demonstrates superior performance over supervised methods.
The paper provides a framework for weakly supervised disentanglement guarantees.
problem Learning disentangled representations in real-world data.
method Theoretical framework for analyzing disentanglement guarantees with weak supervision.
result Empirical verification of weak supervision methods' predictive power and usefulness.
Self-supervised and supervised methods learn similar intermediate visual representations but diverge in final layers.
problem Comparing self-supervised and supervised methods for visual learning.
method Comparison of contrastive self-supervised and supervised methods on simple image data.
result Contrastive and supervised methods learn similar intermediate representations but diverge in final layers.
Semi-supervised clustering methods incorporate a limited amount of supervision into the clustering process. Typically, this supervision is provided by the user in the form of pairwise constraints. Existing methods use such constraints in one of the following ways: they adapt their clustering procedure, their similarity…
Unified framework for classification and image generation.
problem Limited supervision in classification and image generation.
method Triple Generative Adversarial Network (Triple-GAN) as a three-player minimax game.
result Unique equilibrium converges to data distribution, achieving excellent classification and generation results.
The paper tackles semi-supervised learning on point clouds using PDE methods.
problem Extend labels to an entire data set on point clouds.
method Minimizing constrained discrete p-Dirichlet energy, connecting to continuum p-Dirichlet energy, applying PDE methods like pseudo-spectral methods. result Consistency of the numerical scheme in the large data limit for density estimation methods.
The paper improves semi-supervised learning for large data by correcting inconsistent methods.
problem Inconsistent behavior of semi-supervised learning methods in large data limits.
method Data-driven parametrization and theoretical analysis of asymptotic performances.
result Significant performance gains observed on practical data classification.
SS3M learns disease phenotypes from few labels.
problem Lack of supervised data for disease phenotyping.
method Semi-Supervised Mixed Membership Model (SS3M).
result SS3M learns interpretable disease phenotypes.
Improved voice conversion with semi-supervised learning.
problem Voice conversion with limited parallel data.
method Amortized variational inference with parallel and non-parallel utterances.
result Semi-supervised training improves voice conversion performance.
We study high-dimensional asymptotic performance limits of binary supervised classification problems where the class conditional densities are Gaussian with unknown means and covariances and the number of signal dimensions scales faster than the number of labeled training samples. We show that the Bayes error, namely t…
Self-supervision improves GCNs' generalizability and robustness.
problem Improving graph convolutional networks' performance.
method Three mechanisms of self-supervision, multi-task learning, and graph adversarial training.
result Self-supervision enhances GCNs' robustness and generalizability.
The Fredholm integral equation of the first kind improves solutions for ill-posed supervised learning problems with limited data.
problem Ill-posed supervised learning problems with insufficient data.
method Using the Fredholm integral equation of the first kind (FIFK) with semi-supervised assumptions and MSDF methods.
result Improved accuracy and stability in solutions for ill-posed problems.
Paper proposes Universum GANs to improve GANs with limited labeled data.
problem Limited labeled data makes supervised learning challenging.
method Proposes Universum GANs with evolving discriminator loss.
result Improved discriminator accuracy and high quality data generation.
Active WeaSuL uses active learning to improve weak supervision for better model performance.
problem Limited labelled data in machine learning.
method Combines active learning with weak supervision to improve probabilistic labels.
result Active WeaSuL outperforms weak supervision and active learning with limited labelled data.
Self-supervised learning improves representation from EEG signals without labels.
problem Limited supervised data for EEG signal analysis.
method Predicting temporal context from unlabeled EEG time series.
result Self-supervised approach outperforms supervised methods in low data regimes.
Consider a classification problem where we have both labeled and unlabeled data available. We show that for linear classifiers defined by convex margin-based surrogate losses that are decreasing, it is impossible to construct any semi-supervised approach that is able to guarantee an improvement over the supervised clas…
New CNNs learn features from unlabeled data for better activity recognition.
problem Limited labeled data for activity recognition leads to poor generalization.
method Semi-supervised CNNs that learn features from raw sensor data.
result Semi-supervised CNNs outperform supervised and traditional methods by up to 18%.
Paper presents a probabilistic diagnostic model for identifying and treating supervised learning degradation issues.
problem Degradation problems in supervised learning, including class imbalance, overlapping, small-disjuncts, noisy labels, and sparseness.
method Develops a novel probabilistic diagnostic model to identify and treat degradation issues in supervised learning.
result Early and correct diagnosis of degradation issues allows for selecting appropriate remediation treatments and unbiased performance metrics.
A new Lγ-PageRank method enhances semi-supervised learning performance.
problem Limited labeled data and fuzzy graphs hinder classification performance.
method Proposes Lγ-PageRank, a novel approach based on powers of the Laplacian matrix, for signed graphs. result Optimal γ significantly improves classification performance. The paper explains how data augmentation improves semi-supervised learning efficiency.
problem Improving accuracy from a small fraction of labeled data.
method Data augmentation induces a similarity graph, which is graph-Laplacian-regularized for downstream learning.
result A fast transductive rate of O(1/nL) is achieved, reducing the number of labels needed. ProbKT uses probabilistic logical reasoning to train object detection models with weak supervision.
problem Training object detection models requires instance-level annotations, which are often unavailable.
method ProbKT, a framework based on probabilistic logical reasoning, uses arbitrary types of weak supervision.
result ProbKT leads to significant improvement and better generalization compared to existing baselines.
Study shows reverberant phase is not essential for weakly-supervised dereverberation.
problem Evaluating the role of reverberant phase in weakly-supervised dereverberation.
method Statistical Wave Field Theory and recent weak supervision framework.
result Wet phase carries limited useful information and is not essential for weakly supervised dereverberation.
Improves risk control in predictions using semi-supervised calibration.
problem Noisy hyper-parameter tuning from limited labeled data.
method Semi-supervised calibration using unlabeled data to tune hyper-parameters rigorously.
result Improves prediction accuracy without sacrificing statistical validity.
Unified framework for SSL methods linking contrastive and non-contrastive approaches.
problem Lack of theoretical foundations and design guidelines for SSL methods.
method Spectral manifold learning framework to unify SSL methods.
result Theoretical bridge between contrastive and non-contrastive methods.
Proposes a DR method for HSI classification with limited labeled data.
problem High dimensionality, singularity, limited training samples, lack of labeled data, heteroscedasticity, nonlinearity.
method Semi-supervised graph-based, HGF for smoothing, MLP for linear patches, RMLSC and NPMLSC for dissimilarity matrices, optimization for projection matrix.
result Improves HSI classification with limited labeled data.
Self-supervised learning improves EEG signal analysis without labeled data.
problem Limited labeled data in clinical EEG signals.
method Temporal context prediction and contrastive predictive coding tasks.
result SSL-learned features outperform supervised deep neural networks in low-labeled data regimes.
SSNAS finds neural architectures without labeled data.
problem Limited labeled data for NAS.
method Self-supervised learning for NAS.
result Comparable results to supervised NAS with labeled data.
Unified model identifies and locates thoracic abnormalities with limited annotations.
problem Accurate identification and localization of thoracic abnormalities require large annotated datasets, which are expensive to acquire.
method Unified model that simultaneously identifies and localizes abnormalities using limited location annotations.
result Unified model significantly outperforms baseline in classification and localization tasks.
This paper provides an overview of deep semi-supervised learning methods.
problem Reducing the need for large annotated datasets in deep learning.
method Summarizes dominant semi-supervised approaches in deep learning.
result Provides a comprehensive overview of deep semi-supervised learning.