Proposes GM-PLL for better partial label learning.
problem Learning from data with partially labeled instances.
method Reformulates PLL as graph matching problem, incorporating GM scheme and extending matching algorithm.
result Superior performance compared to state-of-the-art methods.
Proposes methods to recover labels from shuffled networks using graph averages.
problem Recovering labels from a shuffled network using graph averages.
method Cluster networks into classes, then match the new graph to cluster-averages, minimizing the graph matching objective function.
result Higher fidelity matching performance when clustering networks into different classes.
A new method for adapting to label shifts using class probability matching.
problem Adapting to label shifts where class probabilities differ between source and target domains.
method Class Probability Matching using Kernel Methods (CPMKM) framework.
result CPMKM outperforms existing methods on real datasets.
Paper quantifies label shift with robustness guarantees using distribution feature matching.
problem Estimating target label distribution under label shift.
method Distribution feature matching (DFM) framework and robustness analysis.
result General performance bound and robustness analysis in misspecified settings.
CM algorithm matches Shannon's and semantic channels for multi-label classification.
problem Tackles label learning and selection for multi-label classification.
method Adheres to maximum semantic information criterion, uses Bayes' theorem, and trains truth functions.
result Shows improved performance and adaptability to changing source distributions.
Bayesian method matches uncertainty to adapt across domains.
problem Label distribution shift across domains degrades model performance.
method Bayesian neural network quantifies uncertainty; joint feature and label distribution matching.
result Improves model performance on domain adaptation tasks.
Study sharpens threshold for matching correlated graphs without labels.
problem Matching latent vertex correspondences in correlated random graphs.
method Analyzes information-theoretic limits for correct vertex matching in sub-sampled graphs.
result Establishes a sharp information-theoretic threshold for vertex matching recovery.
Algorithm aligns correlated Erdős-Rényi graphs efficiently.
problem Graph alignment in correlated Erdős-Rényi graphs.
method A canonical labeling algorithm with two steps: degree-based matching and bipartite graph alignment.
result The algorithm succeeds in aligning correlated Erdős-Rényi graphs in a specific time complexity region.
New research shows semantic data matching can degrade SSDL performance.
problem The limits of semantic data set matching in semi-supervised learning.
method Demonstrated through simulations and a new dissimilarity measure.
result Semantic data matching can degrade SSDL performance under non-IID data.
JPLink uses machine learning to match jobs with RIASEC labels.
problem Matching jobs with RIASEC labels requires significant manual effort.
method JPLink uses text content and O*NET knowledge to assign RIASEC labels to jobs.
result JPLink outperforms conventional baselines in matching jobs with RIASEC labels.
Paper tackles musical version matching at segment level using contrastive learning from weakly-labeled data.
problem Match musical versions at the segment level, not just tracks, with weak annotations.
method Proposes contrastive learning from weakly-labeled audio segments, using a new loss variant.
result Breakthrough performance in segment-level evaluation, outperforming state-of-the-art.
New algorithm improves GAN performance with minimal labels.
problem Improving GAN performance with little supervision.
method Intentionally corrupts generated labels to match real data statistics, trains discriminator with corrupted labels.
result Minimizing proposed loss is equivalent to minimizing true divergence between real and generated data.
ELSA efficiently adapts to label shift without post-prediction calibrations.
problem Domain adaptation with label shift across training and testing datasets.
method Moment-matching framework based on influence function geometry; solves linear systems for adaptation weights.
result ELSA estimator is n \sqrt{n} n -consistent and asymptotically normal, achieving state-of-the-art estimation performance. Most prior work on active learning of classifiers has focused on sequentially selecting one unlabeled example at a time to be labeled in order to reduce the overall labeling effort. In many scenarios, however, it is desirable to label an entire batch of examples at once, for example, when labels can be acquired in para…
Paper proposes a new method for population-wise matching of sulcal graphs.
problem Challenges in matching cortical fold variations across individuals.
method Population-wise multi-graph matching of sulcal graphs.
result Effectiveness of multi-graph matching in obtaining consistent labeling of sulcal basins.
Method addresses label shift in adversarial domain adaptation.
problem Label shift in behavioral studies.
method DATS (Domain Adversarial nets for Target Shift) framework.
result DATS framework performs well under large label shift.
Efficiently matches subgraphs in noisy data without node labels.
problem Subgraph isomorphism in noisy, real-valued graphs.
method Two-step approach: extract topology, then expand matches.
result Realistically sub-linear computational efficiency, robustness to noise.
Prototype Matching Network (PMN) improves genomic TFBS prediction.
problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.
Unified view of label shift estimation methods.
problem Label distribution changes but class-conditional distributions remain the same.
method Unified view of two approaches: BBSE and MLLS.
result Unified framework and theoretical characterization of MLLS.
New method for semi-supervised learning in federated learning with and without labels at clients.
problem Training federated learning models with partially or completely unlabeled data.
method Federated Matching (FedMatch) with inter-client consistency loss and disjoint learning.
result FedMatch outperforms local semi-supervised learning and naive federated learning combinations.
Entity resolution (ER) presents unique challenges for evaluation methodology. While crowdsourcing platforms acquire ground truth, sound approaches to sampling must drive labelling efforts. In ER, extreme class imbalance between matching and non-matching records can lead to enormous labelling requirements when seeking s…
Improves conditional GANs' robustness to noisy labels.
problem Learning conditional generators from noisy labeled samples.
method Introduces Robust Conditional GAN (RCGAN) and RCGAN-U architectures.
result Improves quality and accuracy of generated samples from noisy labels.
A geometric theory explains loss functions for robust representation learning.
problem Treats robustness, domain adaptation, and sensor drift as separate literatures.
method Estimates covariance Sigma_task and uses it to pin Jacobian penalties.
result Proves optimality and necessity of range coverage for penalty matrices.
Optimal transport aligns source and target distributions for domain adaptation.
problem Unsupervised domain adaptation with joint class-conditional and label shifts.
method Minimizes importance weighted loss and Wasserstein distance for aligned marginals and class-conditional distributions.
result Our method outperforms competitors on various domain adaptation tasks.
Method transfers label function spectrum between graphs.
problem Domain adaptation with abrupt label function variations.
method Learning aligned graph bases to transfer label function spectrum.
result Improved classification performance compared to existing methods.
New methods detect targets from imprecisely labeled hyperspectral data.
problem Challenges in acquiring labeled hyperspectral data.
method Multi-Target MI-ACE and MI-SMF methods that learn target signatures from imprecisely labeled samples.
result Effective at learning target signatures and performing target detection.
Crowdsourced labeling recovers task types with minimal queries.
problem Labeling tasks accurately with minimal queries.
method Worker clustering, skill estimation, weighted majority voting.
result Achieves any targeted recovery accuracy with minimum queries.
Retraining with predicted labels improves model accuracy in noisy settings.
problem Improving model accuracy with noisy or corrupted labels.
method Retraining with predicted hard labels in a linearly separable binary classification setting.
result Retraining with predicted labels can increase model accuracy, as proven theoretically.
Paper improves preterm birth prediction using neural networks with noisy labels.
problem Predicting preterm birth from noisy EHR diagnosis codes.
method Developed ALC method to correct label noise in deep learning models.
result Improved prediction performance compared to baseline methods.
Proposes MGPLL for PL learning with non-random noise.
problem Partial label learning with non-random label noise.
method Bi-directional mapping framework, conditional noise label generation, multi-class predictor, adversarial learning.
result Demonstrates state-of-the-art performance in partial label learning.
Efficiently poisons offline RLHF models by flipping preference labels.
problem Vulnerability of offline RLHF models to preference label flipping attacks.
method Developed two attack methods: BAL-A and BMP-A, solving a structured binary sparse approximation problem.
result Demonstrated that flipping one preference label induces a parameter-independent shift in the DPO gradient, enabling structured binary sparse approximation.
Bayesian relabeling improves PU learning for biased positive examples.
problem Learning from biased positive examples in PU learning.
method Probabilistic-gap based PU learning with Bayesian optimal relabeling and kernel mean matching.
result The proposed algorithm improves performance on real-world datasets.
This work analyzes label embedding for large multiclass classification problems.
problem Label embedding for large multiclass classification problems.
method Analysis of label embedding in extreme multiclass classification, presenting an excess risk bound and showing a trade-off between computational and statistical efficiency.
result The statistical penalty for label embedding vanishes with sufficiently low coherence under the Massart noise condition.
New algorithms reduce label collection for online prediction with expert advice.
problem Efficiently predicting binary sequences with expert advice using fewer labels.
method Adaptive selective sampling for exponentially weighted forecasters.
result Label complexity scales roughly as the square root of the number of rounds for a scenario with a strictly better expert.
We analyze decision boundaries using topological data analysis.
problem Quantifying deep neural network complexity for model selection.
method We use labeled Čech complex, plain labeled Vietoris-Rips complex, and locally scaled labeled Vietoris-Rips complex to infer persistent homology of decision boundaries.
result We provide theoretical conditions and analysis for recovering the homology of a decision boundary from samples.
A new model embeds word and label hierarchies in hyperbolic space for HMLC.
problem Learning mappings from word hierarchies to label hierarchies in hierarchical multi-label classification.
method Proposes a Hyperbolic Interaction Model (HyperIM) to learn label-aware document representations in hyperbolic space.
result Demonstrates improved performance for HMLC compared to state-of-the-art methods.
LGGAN generates labeled graphs from graph data.
problem Training generative models for graph-structured data with labels.
method LGGAN, a GAN approach, trains deep models for graph data with node labels.
result LGGAN generates diverse labeled graphs that match training data and outperforms alternatives.
Improved image generation with fewer labels.
problem Generating high-fidelity images with limited labeled data.
method Self- and semi-supervised learning techniques.
result Outperforms state-of-the-art models using 10-20% of labels.
Study shows semi-supervised learning can be more robust with fewer labeled examples.
problem Learning robust predictors in semi-supervised PAC model with minimal labeled data.
method Characterizes the minimal labeled and unlabeled data required for robust learning.
result Proves nearly matching upper and lower bounds on labeled sample complexity.
A blind scheme combines multiple classifiers without knowing their training labels.
problem Combining multiple classifiers to achieve high performance.
method Moment matching method using tensor and matrix factorization.
result Proposed blind scheme outperforms known methods on synthetic and real datasets.
New method improves domain generalization by matching object representations.
problem Existing domain generalization methods fail to generalize to unseen domains.
method Proposes matching-based algorithms to match object representations across domains.
result MatchDG algorithm matches ground-truth object representations and improves out-of-domain accuracy.
New method detects money laundering in Bitcoin using minimal labels.
problem Detecting money laundering in Bitcoin transactions with scarce labels.
method Active learning approach to anomaly detection.
result 5% of labels are sufficient to match supervised baseline performance.
PPI uses predictions and weighting to infer from partially labeled data.
problem Valid inference with partially labeled data.
method Combines model-based predictions with bias correction from labeled data, using Horvitz-Thompson and Hájek corrections.
result IPW-adjusted PPI with estimated propensities performs similarly to known-probability case.
This paper improves multi-label classification by leveraging high-order label correlations.
problem Improving accuracy in multi-label classification tasks using label correlations.
method Exploiting high-order label correlations through a supervised learning classifier system (UCS) and label powerset (LP) strategy.
result The proposed method outperforms other LP-based methods on multiple benchmark datasets.
Automatically evaluates image quality based on human judgment.
problem Difficulty in rigorously evaluating generated image quality.
method Generative model embeddings, human labels regression, and statistical matching.
result 66% accuracy in predicting human scores of image realism.
Linear-time algorithm for optimal assignment in graph matching.
problem Finding optimal assignments between graph vertices efficiently.
method Developed an algorithm for linear-time optimal assignment using tree distances.
result Approximated edit distance between graphs in linear time.
Crowdsourcing systems are popular for solving large-scale labelling tasks with low-paid workers. We study the problem of recovering the true labels from the possibly erroneous crowdsourced labels under the popular Dawid-Skene model. To address this inference problem, several algorithms have recently been proposed, but …
Paper proposes a new model using consistency regularization for learning from label proportions.
problem Learning from label proportions with weak labels on bags of instances.
method Consistency regularization applied to semi-supervised learning.
result LLP with consistency regularization achieves superior performance.