Hashing has been widely used for large-scale approximate nearest neighbor search because of its storage and search efficiency. Recent work has found that deep supervised hashing can significantly outperform non-deep supervised hashing in many applications. However, most existing deep supervised hashing methods adopt a …
This paper provides an overview of deep semi-supervised learning methods.
problem Reducing the need for large annotated datasets in deep learning.
method Summarizes dominant semi-supervised approaches in deep learning.
result Provides a comprehensive overview of deep semi-supervised learning.
S4 learns new self-supervision automatically, improving accuracy with less human effort.
problem Lack of direct supervision in machine learning.
method Combines deep learning and probabilistic logic to automatically generate and verify new self-supervision.
result S4 can automatically propose accurate self-supervision, matching supervised methods with less human effort.
New model handles noisy labels in semi-supervised classification.
problem Noisy class labels in classification tasks.
method M-VAE: A semi-supervised deep generative model that explicitly models noisy labels.
result M-VAE performs better than models ignoring label noise.
Large amounts of labeled data are typically required to train deep learning models. For many real-world problems, however, acquiring additional data can be expensive or even impossible. We present semi-supervised deep kernel learning (SSDKL), a semi-supervised regression model based on minimizing predictive variance in…
Deep nearest neighbors outperform self-supervised methods in anomaly detection.
problem Anomaly detection using self-supervised deep methods.
method Simple nearest-neighbor approach on Imagenet pretrained features.
result Nearest-neighbor method outperforms self-supervised methods in accuracy, few shot generalization, training time, and noise robustness.
Efficiently builds diverse sub-model ensembles for robust self-supervised learning.
problem Challenges in diversity and efficiency of deep ensembles for self-supervised representation learning.
method Ensemble of independent sub-networks with a new loss function for diversity.
result Significantly improves prediction reliability and model calibration.
A simple self-supervised model for tensor RPCA using deep unfolding.
problem Tensor robust principal component analysis (RPCA) challenges in practical applications.
method Deep unfolding with only four hyperparameters.
result Competitive or superior performance compared to supervised methods, even in data-starved scenarios.
Proposes a method to train deep neural networks with limited labeled data.
problem Lack of labeled data in neural text classification.
method Two modules: pseudo-document generator and self-training module.
result Significantly outperforms baseline methods without excessive labeled data.
AutoEmbedder clusters unlabeled data using semi-supervised DNN embedding.
problem Clustering unlabeled data efficiently and effectively.
method Semi-supervised DNN embedding system using Siamese network architecture.
result AutoEmbedder outperforms existing DNN-based semi-supervised methods.
Deep semi-supervised learning identifies tree species from natural images.
problem Identifying tree species in natural settings with limited labeled data.
method Two-fold approach using deep semi-supervised learning.
result Achieves 94.04% top-5 accuracy for leaves and 83.04% for bark.
Unified framework improves deep multi-view clustering by addressing self-supervision and contrastive alignment issues.
problem Variations in self-supervision-based methods for deep multi-view clustering.
method Unified DeepMVC framework that includes recent methods and leverages self-supervision and contrastive alignment.
result Contrastive alignment negatively impacts cluster separability, especially with many views.
Self-supervised learning helps train deep features without needing lots of labeled data.
problem Annotation bottleneck in deep learning.
method Four main families of self-supervised approaches applied to various data modalities.
result Self-supervised methods can now rival fully supervised pre-training across multiple data types.
New method separates objects from images using deep neural networks trained to inpaint.
problem Fully self-supervised instance separation of occluded objects in images.
method Maximizes independence of two image regions given a fully self-supervised inpainting network.
result Method achieves similar segmentation performance to fully supervised methods on microscopy image datasets.
Simplifies transfer learning with deep neural networks using ridge regression.
problem High computational cost of finetuning deep models for transfer learning.
method Leverage the low-rank property of deep neural networks' feature vectors in kernel ridge regression.
result Successful on supervised and semi-supervised transfer learning tasks.
Deep SAD improves anomaly detection with semi-supervised deep learning.
problem Improving anomaly detection with limited labeled data.
method End-to-end deep semi-supervised anomaly detection using information-theoretic framework.
result Deep SAD outperforms shallow, hybrid, and deep competitors on various datasets.
BOSS learns from one labeled sample per class to match fully supervised performance.
problem Achieving fully supervised performance with minimal labeled data.
method Combines class prototype refining, class balancing, and self-training.
result BOSS achieves comparable test accuracies to fully supervised learning.
Two deep learning models improve indoor location prediction from WiFi fingerprints.
problem Indoor location prediction from WiFi fingerprints.
method Convolutional mixture density recurrent neural network and VAE-based semi-supervised learning model.
result Proposed models outperform existing methods in real-world datasets.
Verification determines whether two samples belong to the same class or not, and has important applications such as face and fingerprint verification, where thousands or millions of categories are present but each category has scarce labeled examples, presenting two major challenges for existing deep learning models. W…
Semi-supervised learning algorithms reduce the high cost of acquiring labeled training data by using both labeled and unlabeled data during learning. Deep Convolutional Networks (DCNs) have achieved great success in supervised tasks and as such have been widely employed in the semi-supervised learning. In this paper we…
Study investigates one-shot semi-supervised learning for image classification.
problem Training deep networks requires many labeled samples, limiting adoption.
method Empirical investigation of FixMatch method for one-shot semi-supervised learning.
result Uneven class accuracy is a barrier to high performance in one-shot semi-supervised learning.
Paper proposes semi-supervised learning for EEG analysis.
problem Reducing workload and delays in analyzing large unlabeled EEG datasets.
method Semi-supervised deep learning algorithm using minimal labeled data.
result Predictions can be made with as little as 5 labeled examples.
CycleCluster uses clustering to improve deep semi-supervised learning.
problem Difficulty in obtaining labelled data for deep learning.
method Proposes a new framework using clustering regularisation and graph-based pseudo-labels.
result Demonstrates improved predictive capability through numerical results.
COMBO network improves optical flow estimation by combining deep learning with brightness constancy.
problem Optical flow estimation using deep learning requires complex training schemes.
method COMBO network explicitly exploits brightness constancy and combines it with a data-driven approach.
result COMBO network outperforms state-of-the-art methods on various benchmarks.
Proposes a new method to learn distance metrics for semi-supervised learning.
problem Inconsistency between perturbed input sets and lack of pairwise relationship information.
method Metric Learning by Similarity Network (MLSN) co-training with a classification network to learn distance metrics adaptively.
result Performs better than state-of-the-art methods on empirical tasks.
Recent works demonstrated the usefulness of temporal coherence to regularize supervised training or to learn invariant features with deep architectures. In particular, enforcing smooth output changes while presenting temporally-closed frames from video sequences, proved to be an effective strategy. In this paper we pro…
Boosts few-shot learning with self-supervised representations.
problem Improving generalization on small labeled datasets.
method Introduces self-supervised tasks as auxiliary loss functions.
result Reduces relative error rate by 5-25% on few-shot learning benchmarks.
Paper proposes semi-supervised learning for bearing anomaly detection.
problem Challenges in obtaining accurate labels for bearing fault diagnosis.
method Uses deep variational autoencoders for semi-supervised learning.
result Improves anomaly detection accuracy by 3% to 30% using semi-supervised learning.
Autoencoders improve unsupervised and semi-supervised learning with new generalization bounds.
problem Lack of theoretical understanding of autoencoder generalization in unsupervised and semi-supervised learning.
method Utilized recent advances in deep learning theory and a novel reconstruction loss to provide generalization bounds.
result First theoretical generalization bounds for autoencoders in unsupervised and semi-supervised learning.
Semi-supervised deep learning detects problematic reads for genome assembly.
problem De novo genome assembly is hindered by specific types of reads.
method Analysis of coverage graphs converted to 1D-signals using semi-supervised deep learning models.
result Semi-supervised deep learning models can detect problematic reads with minimal labeled data.
GraphXCOVID uses deep semi-supervised learning to identify COVID-19 from chest X-rays with minimal labels.
problem Identifying COVID-19 on chest X-rays with limited labelled data.
method Graph-based deep semi-supervised framework with pseudo-labeling and attention maps.
result Outperforms current supervised models with a tiny fraction of labelled examples.
Proposes a new framework for semi-supervised learning with theoretical support.
problem Lack of theoretical support for using predictions as pseudo-labels in deep SSL methods.
method D2 framework with repetitive reprediction (R2) strategy.
result R2-D2 method outperforms state-of-the-art methods by 5 percentage points on ImageNet.
Paper tackles domain invariant sentiment classification using weak supervision.
problem Learning a sentiment classification model that adapts to any target domain.
method Two-stage training procedure with weakly supervised datasets.
result Transfer learning with weak supervision achieves performance close to supervised training.
Self-supervised method learns from unlabelled point clouds by reconstructing them.
problem Efficiently learning from large, unlabelled 3D point cloud datasets.
method Trains neural networks to reconstruct point clouds with randomly rearranged parts.
result Method learns semantic properties of point clouds and improves downstream object classification.
DSL uses supervised learning to optimize portfolios, improving stability and performance.
problem Optimizing robust portfolios in financial markets.
method DSL reframes portfolio construction as a supervised learning problem, using cross-entropy loss and optimizing Sharpe or Sortino ratios. Deep Ensemble methods are employed to reduce variance.
result DSL outperforms traditional and machine learning methods, achieving higher median returns and more stable risk-adjusted performance.
Deep GNNs and self-supervision boost graph learning at scale.
problem Efficiently deploying GNNs at large scale remains challenging.
method Two large-scale GNNs: a deep transductive node classifier and a very deep inductive graph regressor.
result Award-level performance on MAG240M and PCQM4M benchmarks.
Proposes a new loss function for deep learning image segmentation.
problem Efficiency and reliance on labeled data in deep learning image segmentation.
method Mumford-Shah functional adapted for deep learning.
result Improves efficiency and effectiveness of deep learning for image segmentation.
Improves deep learning with less labeled data using unsupervised projection.
problem Lack of labeled data for deep learning models.
method Modified unsupervised discriminant projection as a regularization term for semi-supervised learning.
result Proposes an algorithm that enhances classification performance with minimal labeled data.
New method estimates covariance in deep heteroscedastic regression without labels.
problem Estimating covariance in deep heteroscedastic models is challenging due to sample-dependent covariance and lack of ground truth.
method Proposes a self-supervised approach using KL Divergence and 2-Wasserstein distance for covariance estimation and a neighborhood-based heuristic for pseudo labels.
result Demonstrates effective pseudo labels and a computationally cheaper yet accurate deep heteroscedastic regression.
Paper improves deep learning models using contrastive predictive coding for semi-supervised learning.
problem Limited labeled data in semi-supervised learning.
method Contrastive predictive coding technique to improve deep learning models with unlabeled data.
result Proposed cpc-SSL and ccpc-SSL models effectively use unlabeled data, scaling well to large datasets.
Method learns feature maps from deep CNN layers for weakly supervised chest pathology localization.
problem Localization of chest pathologies in X-ray images is challenging due to varying sizes and appearances.
method Class-aware deep multiscale feature learning using intermediate feature maps from CNN layers.
result Improves localization performance of small pathologies like nodules and masses.
Generalized R2R handles non-Gaussian noise for deep network training.
problem Training deep networks from noisy data alone.
method Extending R2R to handle various noise distributions.
result GR2R loss is an unbiased estimator of supervised loss.
Study develops a semi-supervised deep ResNet for Wi-Fi mode detection.
problem Utilizing Wi-Fi signals for multimodal transportation mode detection with limited labeled data.
method Semi-supervised deep residual network (ResNet) framework.
result Framework achieves high prediction accuracy (81.8% for walking, 82.5% for biking, 86.0% for driving).
Deep learning for HJB PDEs using synthetic data and residual minimization.
problem Solving Hamilton-Jacobi-Bellman PDEs for optimal control problems.
method Gradient-augmented synthetic dataset for supervised learning, residual minimization.
result Improves accuracy and efficiency of deep learning for HJB PDEs.
DVT transfers knowledge across domains using semi-supervised deep generative models.
problem Reduces labeling effort in domains with abundant labeled examples.
method Deep Variational Transfer (DVT) using a shared latent Gaussian mixture model.
result Achieves state-of-the-art classification performances across different datasets.
A new model improves semi-supervised learning for unstructured data.
problem Manual labeling of unstructured data is expensive, time-consuming, and prone to errors.
method Batch Semi-Supervised Self-Organizing Map (Batch SS-SOM) integrating deep learning techniques.
result Batch SS-SOM performs well in semi-supervised classification and clustering, even with few labeled samples.
The ever-increasing size of modern data sets combined with the difficulty of obtaining label information has made semi-supervised learning one of the problems of significant practical importance in modern data analysis. We revisit the approach to semi-supervised learning with generative models and develop new models th…
ParsNet tackles weakly supervised data streams with a self-evolving deep neural network.
problem Weakly supervised data streams hinder existing data stream algorithms.
method ParsNet uses a self-labelling strategy with hedge (SLASH) and a closed-loop configuration of generative and discriminative training processes.
result ParsNet outperforms other methods in high-dimensional data streams and infinite delay simulations.