DAL uses disentanglement for automatic labeling in GAN-based active learning.
problem Reducing human labeling in GAN-based active learning.
method DAL leverages disentanglement in InfoGAN to automatically label datapoints, deciding human labeling based on disagreement with InfoGAN labels and label correction.
result DAL achieves better performance than existing GAN-based active learning approaches on image classification tasks.
Improved neural network performance with noisy data.
problem Training neural networks on noisy, automatically annotated data.
method Added a noise layer to a neural network architecture to model and handle noise.
result Improved performance by up to 35% on a low-resource NER task.
Develops STC for sequential data with missing labels.
problem Learning from partially labeled and unsegmented sequential data.
method Introduces Star Temporal Classification (STC) using a star token and GTN framework.
result Recover most of supervised baseline performance with up to 70% missing labels.
Automatically mined rules from dependency parsing help neural models learn from less labeled data.
problem Lack of labeled data for aspect and opinion term extraction.
method Automatically mined rules from dependency parsing, applied to auxiliary data, combined with human-annotated data.
result Neural models achieve better performance than state-of-the-art with mined rules and auxiliary data.
Web browser autofill predicts form field labels for convenience.
problem Predicting form field labels for autofill in web browsers.
method Machine learning solution implemented as a web service using Azure ML Studio.
result Improved form field label prediction for autofill.
Paper uses user engagement signals to automatically label training data for AI assistants.
problem Lack of annotated training data for AI assistants.
method Leverages user engagement signals for unsupervised entity labeling and data augmentation.
result Significant accuracy gains in sequence labeling tasks and user-facing results.
StratPPI improves prediction-powered inference with stratified sampling.
problem Improving statistical estimates with limited human-labeled data.
method Combining small human-labeled data with large automatic-labeled data, stratifying data for tighter confidence intervals.
result StratPPI provides substantially tighter confidence intervals than unstratified approaches.
In conventional supervised pattern recognition tasks, model selection is typically accomplished by minimizing the classification error rate on a set of so-called development data, subject to ground-truth labeling by human experts or some other means. In the context of speech processing systems and other large-scale pra…
CascadeML trains multi-label models without hyperparameter tuning.
problem Improving multi-label classification performance through hyperparameter optimization.
method CascadeML uses a two-phase incremental learning approach to train multi-label neural networks.
result CascadeML outperforms other multi-label classification algorithms without hyperparameter tuning.
Gradual ML improves ALSA with less manual labeling.
problem Lack of high-quality labeled data for ALSA.
method Gradual machine learning for automatic labeling of ALSA tasks.
result Performance surpasses unsupervised and supervised alternatives.
Paper proposes a method to automatically detect drift in machine learning models.
problem Detecting changes in class-label data distributions that affect model predictions.
method Self-evaluating predictive model degradation to detect concept drift.
result Effectiveness in automatically detecting and describing concept drift.
Paper presents a method to train NER models without labelled data using weak supervision.
problem Dealing with NER performance drop in new domains without labelled data.
method Weak supervision through automatic annotation and hidden Markov model integration.
result Improvement of about 7 percentage points in entity-level F1 scores. New approach for active learning in overparameterized models.
problem Efficiently labeling datasets in machine learning.
method MaxiMin Active Learning for nonparametric or overparameterized models.
result Automatically identifies decision boundaries and data clusters.
A new model cleans vocal note event annotations in music.
problem Erroneous labels in music datasets.
method Contrastive learning to automatically create local deformations of likely correct labels.
result Transcription model accuracy improves with the proposed strategy.
S4 learns new self-supervision automatically, improving accuracy with less human effort.
problem Lack of direct supervision in machine learning.
method Combines deep learning and probabilistic logic to automatically generate and verify new self-supervision.
result S4 can automatically propose accurate self-supervision, matching supervised methods with less human effort.
We study active learning where the labeler can not only return incorrect labels but also abstain from labeling. We consider different noise and abstention conditions of the labeler. We propose an algorithm which utilizes abstention responses, and analyze its statistical consistency and query complexity under fairly nat…
Automated labeling of intracranial arteries improves accuracy and efficiency.
problem Challenges in accurately labeling intracranial arteries due to variations and limited datasets.
method Graph Neural Network (GNN) combined with hierarchical refinement for improved accuracy.
result Achieved 97.5% node labeling accuracy on a testing set of 105 scans.
Semi-automatic data annotation helps experts label unlabeled samples based on feature space projection.
problem Laborious manual data annotation for machine learning.
method Interactive semi-automatic approach using feature space projection and semi-supervised learning.
result Reduces user annotation effort and improves classification accuracy.
Over the last couple of years, deep learning and especially convolutional neural networks have become one of the work horses of computer vision. One limiting factor for the applicability of supervised deep learning to more areas is the need for large, manually labeled datasets. In this paper we propose an easy to imple…
New method learns auxiliary labels automatically for improved generalisation.
problem Improving generalisation in supervised learning without additional data.
method Trains two neural networks: label-generation and multi-task networks.
result MAXL outperforms single-task learning on 7 image datasets.
Automatically tunes parameters of rule-based systems using labeled data.
problem Tuning parameters of complex rule-based systems efficiently.
method Structured differential learning for approximate gradient descent.
result Successfully adjusts system values for over 100 parameters.
LeagueAI generates synthetic data for better object detection in video games.
problem Laborious work of gathering large amounts of hand-labeled data for machine vision applications.
method Automatic synthetic dataset generation using game 3D models and background.
result Models trained on synthetic data outperformed those on hand-labeled data in precision and reliability.
With the success of deep learning, recent efforts have been focused on analyzing how learned networks make their classifications. We are interested in analyzing the network output based on the network structure and information flow through the network layers. We contribute an algorithm for 1) analyzing a deep network t…
The paper tackles class imbalance in deep learning models and proposes a method to enhance feature extraction.
problem Class imbalance affects deep learning models, especially in imbalanced settings.
method The paper introduces an extension of deep over-sampling to use automatically-generated abstract-labels for weak-supervision.
result The proposed framework significantly improves image classification benchmarks with imbalanced classes.
Recurrent Neural Networks (RNNs) are extensively used for time-series modeling and prediction. We propose an approach for automatic construction of a binary classifier based on Long Short-Term Memory RNNs (LSTM-RNNs) for detection of a vehicle passage through a checkpoint. As an input to the classifier we use multidime…
Gaussian Processes outperform other models in estimating uncertainty for radiology report observation detection.
problem Uncertainty quantification in automatic data labelling for semi-supervised learning in clinical NLP.
method Investigation of uncertainty estimates from various predictive models using NLPP and MMPCL metrics.
result Gaussian Processes provide superior performance in quantifying uncertainty for radiology report observation detection.
Label noise in adversarial training leads to robust overfitting, explained and mitigated.
problem Label noise in adversarial training causes robust overfitting.
method Proposed a method to automatically calibrate labels.
result Consistent performance improvements across various models and datasets.
New system for automatic music emotion recognition considers multiple emotions simultaneously.
problem Automatic recognition of simultaneous and multiplicity of emotions in music.
method Comparison of multilabel and multiclass machine learning algorithms on the Emotify dataset.
result The Geneva Emotional Music Scale 9 is adopted for multilabel and multiclass classification of music emotions.
Internet of things (IoT) applications have become increasingly popular in recent years, with applications ranging from building energy monitoring to personal health tracking and activity recognition. In order to leverage these data, automatic knowledge extraction - whereby we map from observations to interpretable stat…
This research uses deep learning to automatically classify UN resolutions.
problem Manual labeling of UN documents is too time-consuming.
method Utilizes pre-trained deep learning models without traditional training.
result Shows effectiveness in classifying UN resolutions by SDGs.
The monitoring of sleep patterns without patient's inconvenience or involvement of a medical specialist is a clinical question of significant importance. To this end, we propose an automatic sleep stage monitoring system based on an affordable, unobtrusive, discreet, and long-term wearable in-ear sensor for recording t…
Paper tackles stance detection across domains using adversarial domain adaptation.
problem Stance detection in different domains is costly and tedious.
method Adversarial domain adaptation for stance detection.
result Model effectively transfers knowledge for accurate stance detection across domains.
Paper corrects deep learning for noisy labels.
problem Overfitting to imperfectly labeled data.
method Distribution correction approach to handle noisy inputs.
result Significantly higher accuracy compared to alternative methods.
We investigate the automatic classification of patient discharge notes into standard disease labels. We find that Convolutional Neural Networks with Attention outperform previous algorithms used in this task, and suggest further areas for improvement.
Paper proposes semi-supervised learning for EEG analysis.
problem Reducing workload and delays in analyzing large unlabeled EEG datasets.
method Semi-supervised deep learning algorithm using minimal labeled data.
result Predictions can be made with as little as 5 labeled examples.
ActiveLab improves classifier accuracy with fewer annotations by re-labeling.
problem Imperfect labels from multiple annotators in real-world data.
method ActiveLab automatically decides when to re-label examples for better classifier training.
result ActiveLab trains more accurate classifiers with fewer annotations.
SeqSleepNet tackles automatic sleep staging as a sequence-to-sequence problem.
problem Automatic sleep staging as a sequence-to-sequence classification problem.
method End-to-end hierarchical recurrent neural network (SeqSleepNet) with filterbank and attention-based recurrent layers.
result SeqSleepNet achieves high accuracy (87.1% overall accuracy, 83.3% macro F1-score, 0.815 Cohen's kappa) on a publicly available dataset.
HiGitClass automatically classifies GitHub repositories using keywords.
problem Automatic classification of unlabeled GitHub repositories.
method Keyword-driven hierarchical classification with three modules.
result HiGitClass outperforms existing methods in repository classification.
SAMM monitors model drift in data streams without labels.
problem Detecting concept drift in unsupervised data streams.
method Time and space efficient unsupervised streaming algorithm for drift detection; generates explanations for drift.
result SAMM detects useful anomalous events for fraud detection.
New method debiases selection bias in PU classification with exposure data.
problem Binary classification from positive and unlabeled data with selection bias.
method Automatic Debiased PUE (ADPUE) learning method.
result ADPUE outperforms traditional PU learning methods on various datasets.
ProSelfLC improves robustness of deep neural networks by automatically deciding trust in predictions.
problem Training robust deep neural networks requires addressing issues like label noise and low entropy predictions.
method ProSelfLC progressively increases trust in predicted labels over time, considering entropy and learning time.
result ProSelfLC demonstrates improved robustness in both clean and noisy settings through empirical validation.
Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have considered the multi-label structure in birdsong. We propose to use an ensemble of classifier chain…
PDO optimizes LLM prompts without labels, improving performance.
problem Optimizing prompts for LLMs without access to labeled data.
method Pairwise preference feedback, dueling bandits, Thompson Sampling, mutation.
result PDO identifies stronger prompts than label-free methods.
Method curates cost-effective, high-quality datasets using AI models.
problem Costly manual labeling of datasets.
method Probably Approximately Correct Labels (PACL) method.
result Curates high-quality datasets with low overall labeling error.
AFR simplifies reducing reliance on spurious features, improving model performance.
problem Reducing reliance on spurious features for out-of-distribution generalization.
method Automatic Feature Reweighting (AFR) updates the model with a weighted loss.
result AFR improves model performance on benchmarks with minimal compute.
Paper uses weakly-supervised clustering to automatically create network protocol abstractions.
problem Manual definition of abstraction by domain experts is time-consuming.
method Weakly supervised clustering algorithm for automatic abstraction.
result The method successfully matches the reference abstraction with minimal labeled examples.
Paper tackles partial label learning with self-guided retraining.
problem Dealing with partially labeled examples where each instance has a set of candidate labels.
method Unified formulation with constraints for joint training and pseudo-labeling; maximum infinity norm regularization for automatic differentiation; convex-concave optimization problem; upper-bound surrogate objective function.
result Significantly outperforms state-of-the-art partial label learning approaches.
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
problem Improving AVSR performance with labeled and unlabeled data.
method Semi-supervised method using continuous pseudo-labels generated by the same AVSR model.
result Significant improvements in VSR performance on LRS3 dataset.