Proposes a self-paced multi-label learning method to handle diverse labels efficiently.
problem Learning from multi-label data with a large label space is NP-hard and prone to overfitting.
method Self-paced multi-label learning with diversity (SPMLD) approach, incorporating gradual label inclusion and diversity maintenance.
result The proposed SPMLD framework optimizes a non-convex objective function using block coordinate descent.
iRDM selects unlabeled samples for regression without labels, improving model accuracy.
problem Selecting unlabeled samples for regression without label information.
method Iterative representativeness-diversity maximization (iRDM).
result iRDM significantly outperforms supervised ALR, especially with limited labeled samples.
Self-distillation improves model performance by increasing teacher diversity and smoothing predictions.
problem Improving model generalization and performance through self-distillation.
method Interpreting self-distillation as MAP estimation and proposing instance-specific label smoothing.
result Self-distillation enhances model performance by increasing teacher diversity and smoothing predictions.
Proposes SLCVAE to improve text diversity by self-labeling.
problem KL-Vanishing problem in CVAE for diverse text generation.
method Explicit optimizing objective to guide encoder towards best encoder, using a labeling network.
result Improves text diversity while maintaining comparable accuracy.
Paper tackles multi-source transfer learning with diverse labeling volume and reliability.
problem Challenges in multi-source transfer learning with diverse labeling volume and reliability.
method Combines domain similarity and source reliability through a new transfer learning method, and integrates distribution matching and uncertainty sampling in pool-based active learning.
result Demonstrates superior performance over state-of-the-art transfer learning methods.
This paper proposes new methods for ALR that consider informativeness, representativeness, and diversity.
problem Efficiently label samples for regression models with limited labeled data.
method Integrates informativeness, representativeness, and diversity in pool-based sequential active learning.
result Demonstrates effectiveness of new ALR approaches on 12 datasets.
dHMM improves sequential labeling by encouraging diversity.
problem Improving performance of HMM in real-world sequential labeling tasks.
method dHMM incorporates a diversity-encouraging prior over state-transition probabilities.
result dHMM outperforms state-of-the-art methods on benchmark datasets for PoS tagging and OCR.
Proposes methods to improve wisdom of crowds by considering worker diversity and correlations.
problem Improving wisdom of crowds by considering worker diversity and correlations.
method Proposes inference, learning, and teaching methods considering worker diversity and correlations.
result Proposes methods to improve wisdom of crowds by considering worker diversity and correlations.
New approach corrects image bias without labels.
problem Image search results skew towards majority groups.
method Uses visibly diverse control set to select images.
result Significantly improves visible diversity of results.
A new confidence measure improves self-training in biased data.
problem Improving self-training in biased data.
method Proposes a new confidence measure, T-similarity, based on ensemble diversity of linear classifiers.
result Empirically shows the benefit of T-similarity for pseudo-labeling policies on various datasets.
Two new ALR approaches based on GS reduce labeled samples needed for regression.
problem Need substantial labeled samples for regression models, but unlabeled samples are easy to collect.
method Proposes two new ALR approaches based on greedy sampling (GS) to select beneficial unlabeled samples.
result Extensive experiments on various datasets verified the effectiveness and robustness of the approaches.
Paper proposes a distributed algorithm for multi-label feature selection.
problem Maximizing diversity and quality in non-redundant feature selection.
method Greedy algorithm for distributed optimization of submodular plus diversity functions.
result Achieves constant factor approximation of optimal solution in big data settings.
This work tackles semi-supervised federated learning by reducing model gradient diversity.
problem Improving test accuracy in semi-supervised federated learning with limited labeled data.
method Investigates and compares various design choices including consistency regularization loss, Batch Normalization, and Group Normalization.
result Grouping-based model averaging combined with Group Normalization and consistency regularization loss improves test accuracy.
Proposes a method to increase diversity without sacrificing meritocracy.
problem Systemic bias in datasets affecting diversity and meritocracy.
method Optimally flipping outcome labels and training classification models simultaneously.
result The price of diversity is low and sometimes negative, enhancing diversity without significantly affecting meritocracy.
Meta metric learning improves few-shot learning for diverse domains.
problem Few-shot learning struggles with diverse domains and varying label numbers.
method Task-specific learners with metric learning and a meta learner to discover task-specific metrics.
result Meta metric learning achieves superior performance in diverse multi-domain tasks and flexible label numbers.
AutoWS-Bench-101 evaluates automated weak supervision methods for diverse domains.
problem Limited applicability of weak supervision due to difficulty in designing labeling functions.
method Automates labeling function design using a small set of ground truth labels.
result AutoWS methods often require foundation models to outperform simple few-shot baselines.
The paper tackles fast rates in batch active learning with pool-based data.
problem Reduced adaptivity in batch active learning leads to suboptimal results.
method Proposes a stage-wise greedy algorithm that balances informativeness and diversity.
result The algorithm's excess risk matches minimax rates in standard statistical learning settings.
Proposes Vendi Score for evaluating diversity in ML models.
problem Lack of flexible diversity evaluation metrics in ML.
method Integrates ecological and quantum statistical mechanics concepts to define Vendi Score.
result Vendi Score enables flexible diversity evaluation without requiring a reference dataset.
A new active learning method considers both uncertainty and diversity to minimize labeling and decision costs.
problem Classical AL approaches fail to capture data distribution in unlabeled data, leading to mislabeling of outliers.
method CBAL considers classification uncertainty and instance diversity, using a min-max approach to minimize labeling and decision costs.
result Extensive experiments show CBAL outperforms state-of-the-art AL approaches.
JoCoR improves deep learning with noisy labels by reducing network diversity.
problem Learning with noisy labels in deep learning.
method JoCoR uses two networks to make predictions, calculates a joint loss with Co-Regularization, and updates both networks simultaneously.
result JoCoR outperforms state-of-the-art approaches in learning with noisy labels.
VTAB benchmarks diverse visual tasks to assess representation learning effectiveness.
problem Lack of a unified evaluation for general visual representations.
method Developed VTAB, a benchmark for diverse visual tasks, and evaluated many representation learning algorithms.
result VTAB revealed insights into the effectiveness of various representation learning methods.
New method improves text classification without labeled target data.
problem Improving text classification under domain shift without labeled target data.
method Diversity-based generalization using multi-head attention with diversity constraints.
result Method matches state-of-the-art performance without labeled target data.
This paper tackles unsupervised AL for linear regression.
problem Optimally selecting samples to label from unlabeled data.
method Proposes a novel unsupervised pool-based AL approach considering informativeness, representativeness, and diversity.
result Demonstrated effectiveness on 14 datasets using three different linear regression models.
FLAME auto-labels mobile data efficiently on diverse processors.
problem Accurately and efficiently labeling mobile data with unknown labels on heterogeneous processors.
method Self-adaptive auto-labeling system Flame that schedules and executes workloads on mobile processors.
result Flame achieves high labeling accuracy and performance on heterogeneous mobile processors.
A new method MixGDA combines mixup and gradient-based data augmentation for SSL.
problem Improving semi-supervised learning performance with limited labeled data.
method Gradient-based Data Augmentation (GDA) combined with mixup methods.
result MixGDA achieves state-of-the-art performance in various SSL benchmarks.
Two diversity models improve subset selection for image classification tasks.
problem Data scarcity and high costs in human labeling for supervised learning.
method Facility-Location and Disparity-Min models for training data subset selection and active learning.
result Subset selection improves accuracy by 2-3% with less training data.
M4L-JMF tackles multi-typed objects learning, improving on M3L.
problem Learning from multi-typed objects with diverse features and labels.
method Joint matrix factorization to encode and factorize multi-typed bags and their instances.
result M4L-JMF outperforms existing methods on benchmark datasets.
Unified theory explains diversity in ensemble learning.
problem Explaining diversity in ensemble learning across various scenarios.
method Developed a framework revealing diversity as a hidden dimension in bias-variance decomposition.
result Proved exact bias-variance-diversity decompositions for multiple losses in regression and classification.
NLP tasks are often limited by scarcity of manually annotated data. In social media sentiment analysis and related tasks, researchers have therefore used binarized emoticons and specific hashtags as forms of distant supervision. Our paper shows that by extending the distant supervision to a more diverse set of noisy la…
New methods for handling time-varying label noise in time series classification.
problem Temporal label noise in time series classification tasks.
method Proposed methods to estimate temporal label noise function directly from data.
result Our methods lead to state-of-the-art performance under diverse types of temporal label noise.
This paper improves pool-based sequential active learning for regression.
problem Efficiently selecting unlabeled samples for regression models.
method Proposes three criteria (informativeness, representativeness, diversity) and a new ALR approach using passive sampling.
result The new ALR approach significantly improves model performance across various domains.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
problem Recovering latent actions and environment dynamics from action-free trajectories.
method Assumes distinct policies for each demonstrator, identifies latent transitions and policies via matrix factorization.
result Identifies latent transitions and demonstrator policies up to permutation.
Optimizes forecast accuracy and diversity using multi-task deep learning.
problem Forecasting combinations of time series data.
method Multi-task deep learning architecture that selects and combines forecasting models.
result Enhances point forecast accuracy compared to state-of-the-art methods.
Adaptive model scheduling boosts data labeling efficiency.
problem Efficiently labeling diverse data with limited resources.
method Adaptive Model Scheduling framework using deep reinforcement learning and heuristic algorithms.
result 53% reduction in execution time with no loss of labels.
A new method for synthetic oversampling of multi-label data focusing on local label distribution.
problem Class imbalance in multi-label datasets affects prediction accuracy.
method Proposes a new method for synthetic oversampling of multi-label data focusing on local label distribution.
result Demonstrates effectiveness in generating more diverse and better labeled instances.
New method reduces overfitting in deep neural networks by measuring and regulating hidden unit diversity.
problem Overfitting in deep neural networks.
method Introduces a new redundancy measure based on mutual information to improve generalization.
result Reduction of redundancy improves generalization capacity, reducing overfitting.
Proposes methods to improve multi-label learning by addressing local label imbalance.
problem Local label imbalance within minority class examples degrades multi-label learning performance.
method Introduces a measure to assess local label imbalance and two sampling approaches (MLSOL, MLUL) to address it.
result Experimental results show MLSOL and MLUL improve performance on multi-label datasets.
New method selects diverse mini-batches for active learning.
problem Reduce labeled data for deep learning models.
method Sequential selection of diverse mini-batches using K-means clustering.
result Achieves comparable or better performance than previous methods.
Labels distilled from images improve model training efficiency and flexibility.
problem Creating synthetic labels for a small set of real images to train models effectively.
method Introduce a more robust and flexible meta-learning algorithm for distillation and an effective first-order strategy based on convex optimization layers.
result Label distillation leads to improved results and greater flexibility in neural architectures.
A new method combines experts' opinions to train regression models with noisy labels.
problem Training regression models with noisy labels from multiple experts.
method Estimate each labeler's expertise and combine opinions using learned weights.
result Empirically outperforms existing techniques on simulated and real data.
To cope with the high level of ambiguity faced in domains such as Computer Vision or Natural Language processing, robust prediction methods often search for a diverse set of high-quality candidate solutions or proposals. In structured prediction problems, this becomes a daunting task, as the solution space (image label…
New active learning methods use statistical leverage scores to select examples efficiently.
problem Efficiently selecting labeled examples for high model accuracy with limited labeled data.
method Proposes ALEVS and DBALEVS methods based on statistical leverage scores.
result DBALEVS selects diverse, representative examples efficiently.
Unified approach for learning with weak labels across various tasks.
problem Learning with noisy or incomplete labels in diverse machine learning settings.
method Implicit posterior models for joint label inference.
result Unified training objective for various machine learning tasks.
LGGAN generates labeled graphs from graph data.
problem Training generative models for graph-structured data with labels.
method LGGAN, a GAN approach, trains deep models for graph data with node labels.
result LGGAN generates diverse labeled graphs that match training data and outperforms alternatives.
Improved image generation with fewer labels.
problem Generating high-fidelity images with limited labeled data.
method Self- and semi-supervised learning techniques.
result Outperforms state-of-the-art models using 10-20% of labels.
Estimates calibration error under label shift without labels.
problem Ensuring model reliability in the face of dataset shift without access to labels.
method Importance re-weighting of the labeled source distribution to estimate calibration error under label shift.
result Effective and reliable CE estimation with respect to the shifted target distribution.
Proposes ML-GCN for multi-label network node representation learning.
problem Complex multi-label networks with correlated labels.
method Two Siamese GCNs model node-label and label-label interactions, integrated under a unified objective function.
result Effective node representation learning with preserved label interactions.
CCVAE captures label characteristics in VAEs for better representation learning.
problem Capturing rich label characteristics in VAEs without conflating them with label values.
method Developed CCVAE, a novel VAE model that explicitly captures label characteristics in latent space.
result CCVAE allows for effective and general interventions like smooth traversals and diverse conditional generation.