This paper compares unstructured and structured EM-based semi-supervised learning methods.
problem Semi-supervised learning with EM algorithm for structured prediction.
method Comparative study between unstructured and structured EM-based semi-supervised learning methods.
result Structured EM is more robust to class confusion in flood mapping datasets.
Unified approach combines prediction-powered inference and variance reduction for semi-supervised optimization.
problem Scarcity of labeled data in semi-supervised optimization.
method PPI-SVRG, combining PPI and SVRG methods.
result Unified convergence bound with improved performance under label scarcity.
ASGN uses active semi-supervised learning to predict molecular properties efficiently.
problem Predicting molecular properties with scarce labeled data and high computational cost.
method ASGN combines a teacher-student framework with active learning to handle joint representation and property learning.
result ASGN achieves remarkable performance in property prediction on public datasets.
Unified model combines feature and label propagation for semi-supervised classification.
problem Combining feature and label propagation for effective semi-supervised classification.
method Unified Message Passing Model (UniMP) using Graph Transformer and masked label prediction.
result Obtains new state-of-the-art results in Open Graph Benchmark (OGB).
Proposes a probabilistic approach to semi-supervised learning using normalizing flows.
problem Leveraging unlabelled data for semi-supervised learning with limited labelled data.
method Uses a normalizing flow to learn the posterior distribution over predictions for labelled data, serving as a prior for unlabelled data.
result Demonstrates improved performance on various tasks with varying output complexity.
Improves risk control in predictions using semi-supervised calibration.
problem Noisy hyper-parameter tuning from limited labeled data.
method Semi-supervised calibration using unlabeled data to tune hyper-parameters rigorously.
result Improves prediction accuracy without sacrificing statistical validity.
Two deep learning models improve indoor location prediction from WiFi fingerprints.
problem Indoor location prediction from WiFi fingerprints.
method Convolutional mixture density recurrent neural network and VAE-based semi-supervised learning model.
result Proposed models outperform existing methods in real-world datasets.
Large amounts of labeled data are typically required to train deep learning models. For many real-world problems, however, acquiring additional data can be expensive or even impossible. We present semi-supervised deep kernel learning (SSDKL), a semi-supervised regression model based on minimizing predictive variance in…
Use of computational methods to predict gene regulatory networks (GRNs) from gene expression data is a challenging task. Many studies have been conducted using unsupervised methods to fulfill the task; however, such methods usually yield low prediction accuracies due to the lack of training data. In this article, we pr…
Semi-supervised GANs improve travel mode inference from GPS data.
problem Travel mode inference from GPS trajectories.
method Developed semi-supervised GANs and compared them with CNNs on a large-scale smartphone dataset.
result Best semi-supervised GAN model achieved 83.4% prediction accuracy.
Proposes a new framework for semi-supervised learning with theoretical support.
problem Lack of theoretical support for using predictions as pseudo-labels in deep SSL methods.
method D2 framework with repetitive reprediction (R2) strategy.
result R2-D2 method outperforms state-of-the-art methods by 5 percentage points on ImageNet.
New approach combines semi-supervised learning and bandits for better predictions.
problem Online semi-supervised learning with bandit feedback for applications like clinical trials and ad recommendations.
method Adjusted Graph Convolutional Network (GCN) for contextual bandits, with semi-supervised missing rewards imputation.
result Developed multi-GCN embedded contextual bandit algorithms verified on real-world datasets.
ICT improves semi-supervised learning by making predictions consistent at data interpolations.
problem Improving semi-supervised learning performance with limited labeled data.
method ICT encourages consistent predictions at data interpolations, reducing overfitting.
result ICT achieves state-of-the-art performance on CIFAR-10 and SVHN datasets.
SADA safely combines predictions from various models for semi-supervised learning.
problem Combining uncertain quality predictions from multiple models in semi-supervised learning.
method Safe and adaptive aggregation of black-box predictions.
result The method guarantees better performance than using labeled data alone and adapts to perfect predictions.
Better use of unlabelled data improves Bayesian active learning models.
problem Neglecting unlabelled data harms predictive performance and data acquisition decisions.
method A simple framework for semi-supervised Bayesian active learning.
result The proposed framework produces better models than conventional approaches.
Paper improves deep learning models using contrastive predictive coding for semi-supervised learning.
problem Limited labeled data in semi-supervised learning.
method Contrastive predictive coding technique to improve deep learning models with unlabeled data.
result Proposed cpc-SSL and ccpc-SSL models effectively use unlabeled data, scaling well to large datasets.
New framework improves generative models with prediction and consistency constraints.
problem Improving generative models with sparse labeled data.
method Optimizes variational autoencoders with prediction and consistency constraints.
result Promising image classification performance, especially in semi-supervised scenarios.
Semi-supervised GAN creates synthetic genetic data for disease prediction.
problem Expensive and time-consuming to build large labeled genetic databases.
method Semi-supervised Genetic Generative Adversarial Network (gGAN).
result Model achieved satisfactory results with real genetic data.
Supervisory signals have the potential to make low-dimensional data representations, like those learned by mixture and topic models, more interpretable and useful. We propose a framework for training latent variable models that explicitly balances two goals: recovery of faithful generative explanations of high-dimensio…
FixMatch combines consistency and confidence to simplify semi-supervised learning.
problem Improving model performance with unlabeled data.
method Generates pseudo-labels using weak augmentation, retains high-confidence predictions, trains on strongly augmented images.
result Achieves state-of-the-art performance on semi-supervised learning benchmarks.
ModSSC unifies semi-supervised classification for various data types.
problem Fragmented support for semi-supervised classification across different methods, settings, and data types.
method ModSSC is a modular Python framework that supports reproducible and controlled experimentation for semi-supervised classification on heterogeneous data.
result ModSSC enables systematic comparison of semi-supervised learning across various datasets and model backbones.
Study shows semi-supervised learning improves human activity recognition with minimal user input.
problem Improving human activity recognition models using incremental learning.
method Three approaches: non-supervised, semi-supervised, and supervised learning were compared.
result Semi-supervised learning achieves similar accuracy to supervised learning with minimal user input.
A framework uses a mixture of predictors for semi-supervised inference.
problem Limited labeled data, abundant unlabeled data.
method Mixture of Experts (MOE) for semi-supervised inference.
result MOE-powered inference framework achieves smallest possible variance.
Use network embeddings to correct for unobserved confounding.
problem Causal inference in the presence of unobserved confounding.
method Use network embeddings to semi-supervised predict treatments and outcomes.
result Valid causal inferences under suitable conditions on predictive model quality.
MEC improves efficiency and robustness in semi-supervised inference.
problem Efficient inference with limited labeled data and robust uncertainty quantification.
method Machine-Learning-Assisted Generalized Entropy Calibration (MEC) using cross-fitted, calibration-weighted PPI.
result MEC achieves semiparametric efficiency bounds under weaker assumptions and provides near-nominal coverage.
We train and validate a semi-supervised, multi-task LSTM on 57,675 person-weeks of data from off-the-shelf wearable heart rate sensors, showing high accuracy at detecting multiple medical conditions, including diabetes (0.8451), high cholesterol (0.7441), high blood pressure (0.8086), and sleep apnea (0.8298). We compa…
Automates MIPs solution with semi-supervised graph neural networks.
problem Solving recurrent Mixed-Integer Programming (MIP) problems efficiently.
method Semi-supervised Graph Neural Networks (GNNs) for predicting variable values.
result GNNs can solve MIPs with unlabeled data and improve over other ML approaches.
Semi-supervised learning improves prediction using unlabeled data.
problem Improving prediction performance using unlabeled data.
method General methodology for semi-supervised Empirical Risk Minimization (ERM) focusing on generalized linear regression.
result Adaptive SSL can achieve substantial improvement over supervised and null models in various settings.
Deep generative models (DGMs) are effective on learning multilayered representations of complex data and performing inference of input data by exploring the generative ability. However, it is relatively insufficient to empower the discriminative ability of DGMs on making accurate predictions. This paper presents max-ma…
Paper proposes semi-supervised learning for EEG analysis.
problem Reducing workload and delays in analyzing large unlabeled EEG datasets.
method Semi-supervised deep learning algorithm using minimal labeled data.
result Predictions can be made with as little as 5 labeled examples.
Protein function prediction is the important problem in modern biology. In this paper, the un-normalized, symmetric normalized, and random walk graph Laplacian based semi-supervised learning methods will be applied to the integrated network combined from multiple networks to predict the functions of all yeast proteins …
Develops a semi-supervised learning method using exponential tilt mixture models.
problem Improves classification accuracy with labeled and unlabeled data.
method Extends logistic regression to exponential tilt modeling, derives maximum likelihood estimation, and proposes regularized estimation.
result Demonstrates improved prediction accuracy compared to existing methods.
OTI extends OTP for inductive semi-supervised learning.
problem Inductive semi-supervised learning for out-of-sample data.
method Optimal transport-based approach extended to inductive tasks.
result OTI outperforms state-of-the-art methods in experiments.
Semi-supervised GANs with log-signatures improve credit card fraud detection.
problem Detecting fraud in large, complex financial transaction data streams.
method Conditional GANs with Bayesian inference and log-signatures for robust feature encoding.
result Consistent improvements over benchmarks in global and domain-specific metrics.
Local Clustering improves semi-supervised learning models.
problem Improving semi-supervised learning models with limited labeled data.
method Local Clustering (LC) method to mitigate confirmation bias in Mean Teacher (MT) model.
result Adding LC loss to MT improves model performance on semi-supervised benchmark datasets.
S2MAM improves semi-supervised learning by selecting relevant variables and updating similarity metrics.
problem Joint learning from labeled and unlabeled data with geometric structure.
method Bilevel optimization scheme for automatic variable selection and similarity matrix update.
result The proposed S2MAM achieves robust and interpretable predictions.
Semi-supervised learning improves QSAR model predictions for novel compounds.
problem Improving model predictions for compounds not in the training set and adjusting for selection bias.
method Semi-supervised learning framework to estimate model quality and adjust for selection bias.
result Predictions for novel compounds are improved by accounting for compound similarity and selection bias.
New method improves semi-supervised learning with missing labels.
problem Reliable classification with missing labels in semi-supervised learning.
method Develops a new semi-supervised learning approach that relaxes assumptions about unlabeled data.
result Provides classifiers that reliably quantify label uncertainty.
Study of multi-task semi-supervised learning in high dimensions.
problem Analysis of multi-task semi-supervised learning in high-dimensional settings.
method Random matrix theory applied to characterize asymptotics of key functionals.
result Predicts algorithm performance and provides efficient usage guidelines.
Paper proposes semi-supervised method for dictionary learning.
problem Learning from both labeled and unlabeled data.
method Uses semi-supervised dictionary learning with LLE for manifold preservation.
result Significant improvements over other methods demonstrated.
Expanding on SSL, this study considers predicting outcomes from both causes and effects.
problem Understanding the limits and possibilities of semi-supervised learning.
method Explores the conditional distribution of effect features given causal features for semi-supervised learning.
result Shows that semi-supervised learning can be effective when predicting outcomes from both causes and effects.
We propose a simple discrete time semi-supervised graph embedding approach to link prediction in dynamic networks. The learned embedding reflects information from both the temporal and cross-sectional network structures, which is performed by defining the loss function as a weighted sum of the supervised loss from past…
ConfGCN estimates labels and confidences in graph-based semi-supervised learning.
problem Predicting node properties in graphs with limited labeled data.
method ConfGCN uses graph convolutional networks to estimate labels and confidences jointly, improving upon anisotropic neighborhood aggregation.
result ConfGCN outperforms state-of-the-art baselines on standard benchmarks.
FlowGMM uses normalizing flows for semi-supervised learning, showing promising results across various data types.
problem Semi-supervised learning with limited labeled data.
method Normalizing flows combined with latent Gaussian mixture models for generative modeling.
result FlowGMM achieves promising results on multiple data types, including text and tabular data.
NETpred uses graph models to predict multiple market indices.
problem Predicting multiple market indices with high accuracy.
method NETpred constructs a heterogeneous graph of related indices and stocks, selects representative nodes, and uses semi-supervised learning to predict index labels.
result NETpred outperforms state-of-the-art methods by 3%-5% in F-score on various datasets.
Develops a SAS approach for high-dimensional risk prediction using unlabeled data.
problem Challenges in risk modeling with EHR data due to lack of direct disease outcomes and high dimensionality.
method Surrogate Assisted Semi-supervised Learning (SAS) approach leveraging unlabeled and labeled data.
result Valid inference for predicted risk even when underlying model is dense and mis-specified.
Enhances U-statistics for semi-supervised datasets using unlabeled data.
problem Efficiently utilizing unlabeled data in semi-supervised settings.
method Semi-supervised U-statistics enhanced by unlabeled data.
result Proposed method is asymptotically Normal and more efficient than classical U-statistics.
Generative adversarial networks improved for regression tasks.
problem Improving neural network training with semi-supervised data.
method Feature contrasting loss function for discriminator training.
result Semi-supervised GANs can be effectively applied to regression problems.