Paper proposes a weak supervision technique for CNN semantic segmentation of lung diseases using partially annotated data.
problem Creating annotated datasets for semantic segmentation of lung diseases is laborious and time-consuming.
method Proposes a weak supervision technique that utilizes partially annotated datasets to improve CNN semantic segmentation accuracy.
result Significantly improved segmentation accuracy using partially annotated datasets.
Active learning selects both observations and annotation precision for Gaussian Processes.
problem Costly annotation in supervised learning.
method Proposes an active learning algorithm that selects observations and annotation precision, using a modified BALD objective.
result Empirically shows the benefits of adjusting annotation precision in active learning.
Regression problems assume every instance is annotated (labeled) with a real value, a form of annotation we call \emph{strong guidance}. In order for these annotations to be accurate, they must be the result of a precise experiment or measurement. However, in some cases additional \emph{weak guidance} might be given by…
Paper proposes an algorithm to recover full supervision from weakly labeled data.
problem Machine learning requires expensive data annotation, motivating the use of weak supervision.
method The paper introduces a disambiguation principle and an empirical disambiguation algorithm for partial labelling.
result The algorithm achieves exponential convergence rates under learnability assumptions.
Paper presents a method to train NER models without labelled data using weak supervision.
problem Dealing with NER performance drop in new domains without labelled data.
method Weak supervision through automatic annotation and hidden Markov model integration.
result Improvement of about 7 percentage points in entity-level F1 scores. ProbKT uses probabilistic logical reasoning to train object detection models with weak supervision.
problem Training object detection models requires instance-level annotations, which are often unavailable.
method ProbKT, a framework based on probabilistic logical reasoning, uses arbitrary types of weak supervision.
result ProbKT leads to significant improvement and better generalization compared to existing baselines.
Active learning improves inspection systems by using weakly labeled data.
problem Rapidly updating machine vision inspection systems in evolving manufacturing processes.
method Developed a methodology for active learning from weakly labeled data, addressing covariate shift with domain-adversarial training.
result Demonstrated that active learning can accelerate the annotation process and reduce false positives.
Framework detects fake news using weak social signals from multiple sources.
problem Lack of annotated data for early fake news detection.
method Jointly uses weak social signals and clean data to train deep neural networks in a meta-learning framework.
result Framework outperforms state-of-the-art baselines for early fake news detection.
SPIN converts weak LLMs to strong ones using self-play.
problem Growing strong LLMs without human-annotated data.
method Self-play fine-tuning (SPIN) starting from a supervised fine-tuned model.
result SPIN significantly improves LLM performance across benchmarks.
Interactive weak supervision learns useful heuristics from user feedback.
problem Creating useful heuristics for large labeled datasets is tedious and subjective.
method Develops an interactive framework for learning heuristics from user feedback.
result Only a few feedback iterations are needed to train models without ground truth labels.
A method solves weakly supervised multiclass MIL problems.
problem Weakly supervised multiclass MIL problems in image recognition.
method Maximizing exact likelihood fitting given observations.
result Our method can learn all convolutional neural networks.
Machine learning approaches hold great potential for the automated detection of lung nodules in chest radiographs, but training the algorithms requires vary large amounts of manually annotated images, which are difficult to obtain. Weak labels indicating whether a radiograph is likely to contain pulmonary nodules are t…
A new method for weakly supervised learning that improves model accuracy.
problem Training machine learning models with precise labels is expensive; weak supervision provides a low-cost alternative.
method Data consistent weak supervision algorithm that searches over classifiers to find plausible labelings, considering features of the training data and estimating labels for low/no coverage data.
result Empirically, the method significantly outperforms state-of-the-art weak supervision methods on text and image classification tasks.
Automatically mined rules from dependency parsing help neural models learn from less labeled data.
problem Lack of labeled data for aspect and opinion term extraction.
method Automatically mined rules from dependency parsing, applied to auxiliary data, combined with human-annotated data.
result Neural models achieve better performance than state-of-the-art with mined rules and auxiliary data.
Unified framework for structured prediction with partial labelling.
problem Learning with partial labelling costs less but is not well studied.
method Structured prediction and infimum loss for a wide range of problems.
result Unified framework leads to explicit algorithms with statistical consistency.
Detect objects from motion without annotations.
problem Weakly supervised object detection.
method Train model on videos of moving objects and negative scenes.
result Detects objects in single images without annotations.
Convolutional LSTM detects emphysema in lung cancer screening images.
problem Learning disease signatures from weakly annotated volumetric medical images.
method 3D volumetric images analyzed as a sequence of 2D images using convolutional LSTM.
result Convolutional LSTM model outperformed other methods in detecting emphysema.
Faster weak supervision framework using triplet methods.
problem Computational inefficiency in weak supervision models.
method Closed-form solution for latent variable models, avoiding iterative methods.
result Orders of magnitude faster than previous approaches.
The scarcity of data annotated at the desired level of granularity is a recurring issue in many applications. Significant amounts of effort have been devoted to developing weakly supervised methods tailored to each individual setting, which are often carefully designed to take advantage of the particular properties of …
Improves data labeling efficiency in machine learning.
problem Efficiency in data labeling for machine learning models.
method Weakly supervised learning, active labeling, stochastic gradient descent.
result Derives a new algorithm for active labeling that scales better with input dimension.
In this paper, we propose a method for training neural networks when we have a large set of data with weak labels and a small amount of data with true labels. In our proposed model, we train two neural networks: a target network, the learner and a confidence network, the meta-learner. The target network is optimized to…
WSGN detects actions from weak supervision, improving performance on THUMOS14 and Charades.
problem Challenging action detection requires detailed manual supervision.
method WSGN learns action detection from video-level labels, exploiting both video-specific and dataset-wide statistics.
result WSGN achieves significant gains in action detection for THUMOS14 and Charades datasets.
Bayesian U-Net exploits epistemic uncertainty for anomaly detection in retinal OCT images.
problem Anomaly detection in retinal OCT images using weak labels of healthy anatomy.
method Bayesian U-Net trained on weak labels of healthy anatomy, using Monte Carlo dropout for uncertainty estimation, and post-processing to transfer uncertainty to anomaly segmentations.
result Achieved a Dice index of 0.789 in an independent test set of AMD cases.
Paper proposes a new model using consistency regularization for learning from label proportions.
problem Learning from label proportions with weak labels on bags of instances.
method Consistency regularization applied to semi-supervised learning.
result LLP with consistency regularization achieves superior performance.
Training deep neural networks requires massive amounts of training data, but for many tasks only limited labeled data is available. This makes weak supervision attractive, using weak or noisy signals like the output of heuristic methods or user click-through data for training. In a semi-supervised setting, we can use a…
Study shows annotation instrument design affects model performance in hate speech detection.
problem Impact of annotation instrument design on model performance in hate speech detection.
method Collected annotations from five experimental conditions of an annotation instrument, fine-tuned BERT models on each dataset, evaluated performance on holdout portion.
result Significant differences in model performance and annotations across conditions.
Paper tackles brain tumor segmentation using weak labels and hierarchical training.
problem High manual labeling effort for fully supervised brain tumor segmentation.
method Uses scribbles and global labels for weak supervision, trains two networks in phases.
result Achieves competitive results on brain tumor segmentation and substructure segmentation.
New method to estimate doctors' effort in annotating medical images.
problem High effort and expense in annotating medical images.
method Proposes a new criterion to evaluate effort, uses active learning and U-shape network for annotation strategy, and fine annotation platform to reduce effort.
result State-of-the-art segmentation performance achieved with only 60% annotation candidates, reducing effort by 44-47%.
Algorithm selects best sensor tests for unknown outcomes.
problem Selecting optimal sensors in unsupervised systems.
method Developed an algorithm for stochastic partial monitoring under Weak Dominance property.
result Algorithm achieves sub-linear regret in sensor selection.
One of the problems on the way to successful implementation of neural networks is the quality of annotation. For instance, different annotators can annotate images in a different way and very often their decisions do not match exactly and in extreme cases are even mutually exclusive which results in noisy annotations a…
Generative model combines multi-dimensional annotations for more accurate ground truth estimation.
problem Inaccurate ground truth estimation from naive annotators' multi-dimensional annotations.
method Proposes a joint multi-dimensional model for global and time-series annotation fusion using Expectation-Maximization algorithm.
result More accurate ground truth estimates through joint modeling of multiple dimensions.
Survey on AL strategies for cost-effective annotation in classification.
problem Real-world AL challenges due to human annotators' limitations.
method Categorizes 60 real-world AL strategies considering multiple annotators, query types, and cost schemes.
result General real-world AL strategy introduced for categorization of 60 strategies.
RAD improves robustness to domain annotation noise without explicit domain annotations.
problem Robustness to domain annotation noise in training data.
method Regularized Annotation of Domains (RAD) for last layer retraining.
result RAD outperforms state-of-the-art methods even with 5% noise in training data.
Generative-discriminative method improves label prediction and instance generation.
problem Difficulty in obtaining high-quality labeled instances.
method Generative-discriminative complementary learning method using CC-GAN.
result Improves accuracy in predicting ordinary labels and generating high-quality instances.
The study challenges the notion that partial data annotation is inferior, suggesting it can sometimes outperform complete annotation.
problem The inefficiency and high cost of completely annotating structured data.
method Information theoretic formulation applied to three diverse structured learning tasks.
result Learning from partial structures can sometimes outperform learning from complete ones.
New method learns useful disentangled representations from weakly labeled data.
problem Learning useful representations from weakly labeled data.
method Model pairs of non-i.i.d. images, learn disentangled representations without requiring annotation.
result Learn disentangled representations reliably from pairs of images without requiring group, individual factor, or number of changed factors annotation.
A new method uses triplet embeddings to improve human annotation for hidden constructs.
problem Improving human annotation for hidden constructs in machine learning.
method Proposes a novel annotation approach using triplet embeddings to lift absolute annotations to relative comparisons.
result Successfully represents synthetic hidden constructs in time under noisy sampling conditions.
Approach for training deep nets with unlabeled patches.
problem Training deep neural networks with detailed expert annotations.
method Cluster-based learning from weakly labeled bags in latent space.
result Improved performance on Camelyon dataset.
PTBCC improves accuracy in multi-class annotation aggregation by learning from prototype confusion matrices.
problem Inaccurate and insufficient confusion matrices for annotators in multi-class classification tasks.
method PTBCC (ProtoType learning-driven Bayesian Classifier Combination) uses prototype confusion matrices to capture annotator expertise.
result PTBCC achieves up to 15% accuracy improvement and 3% higher average accuracy compared to existing methods.
ActiveLab improves classifier accuracy with fewer annotations by re-labeling.
problem Imperfect labels from multiple annotators in real-world data.
method ActiveLab automatically decides when to re-label examples for better classifier training.
result ActiveLab trains more accurate classifiers with fewer annotations.
Bayesian methods improve text annotation quality.
problem Inconsistent and unreliable human annotations in natural language processing.
method Two semi-supervised Bayesian methods: a deep learning model and an ensemble method.
result Bayesian methods enhance the reliability and performance of BERT models.
Study improves app feature extraction models with new annotation guidelines and data.
problem Improving the quality and usefulness of app feature extraction models.
method Exploring the effects of annotation guidelines and annotated data on app feature extraction models.
result New annotation guidelines lead to less noisy and more informative app features.
Paper proposes an efficient method for bounding box annotation in object detection.
problem Manual annotation of bounding boxes is tedious and resource-intensive.
method Iterative training of object detector on small batches of labeled images, with human annotator correcting errors.
result Significant reduction in human annotation effort, up to 75%.
Paper introduces a semi-supervised method for image segmentation using a small set of labeled images.
problem Costly and time-consuming to create large labeled datasets for semantic segmentation.
method Uses a small set of fully labeled images and a weak set of only bounding box labeled images. Trains a primary model with an ancillary model generating initial labels and a self-correction module improving these labels.
result Models trained with a small fully supervised set perform similarly or better than those trained with a large fully supervised set, requiring significantly less annotation effort.
Improves off-policy evaluation with imperfect annotations.
problem Limited dataset coverage for evaluating new policies.
method Doubly robust estimators combining IS and DM, incorporating counterfactual annotations.
result Using annotations within the DM component yields the most desirable theoretical results.
A popular approach for large scale data annotation tasks is crowdsourcing, wherein each data point is labeled by multiple noisy annotators. We consider the problem of inferring ground truth from noisy ordinal labels obtained from multiple annotators of varying and unknown expertise levels. Annotation models for ordinal…
Crowdlab uses classifiers to estimate consensus labels and annotator quality.
problem Leveraging multiple annotators for data classification.
method Weighted ensemble approach using any trained classifier.
result Superior estimates for consensus labels and annotator quality.
Method learns true labels from noisy annotators using regularization.
problem Learning from noisy labels in supervised learning.
method Regularized estimation of annotator confusion matrices.
result Method outperforms state-of-the-art methods in image classification.