Improves writer identification with unlabeled data and weighted label smoothing.
problem Offline writer identification requires labeled data, which is costly and time-consuming.
method Proposed a semi-supervised feature learning pipeline with weighted label smoothing regularization.
result Significantly improved baseline performance on writer identification datasets.
Paper proposes a new PLL framework with a progressive identification algorithm.
problem Weakly supervised learning with partial labels.
method Flexible model and optimization algorithm for PLL, progressive identification algorithm.
result Established an estimation error bound and set new state of the art.
Unified method for learning from selectively labeled data.
problem Classification with selectively labeled data from multiple decision-makers.
method Unified cost-sensitive learning (UCL) approach.
result Unified method for robust classification in selective labeling.
Can we identify node labels from graph labels?
problem Identifying node labels from graph labels in a hierarchical network.
method Gaussian Mixture Graph Convolutional Network (GMGCN) with Graph Attention Network (GAT) and Gaussian Mixture Layer (GML).
result The proposed method outperforms other baselines on various benchmarks.
New method estimates model performance bounds without ground truth labels.
problem Evaluation of weakly supervised models without direct access to ground truth labels.
method Formulates model evaluation as a partial identification problem and uses Fréchet bounds for performance estimation.
result Derives accurate and computationally efficient bounds for key metrics like accuracy, precision, recall, and F1-score.
Paper improves gas species identification in complex mixtures using neural networks.
problem Identifying gas species in multi-gas mixtures with high accuracy.
method Multi-label neural networks with optimal thresholding for IR spectroscopy.
result Optimal thresholding improves classification performance over conventional methods.
Transfer learning improves NER performance on limited labeled data.
problem Label scarcity in patient note de-identification.
method Transfer learning from a large labeled dataset to a smaller labeled dataset.
result Transfer learning enhances NER performance on patient note de-identification.
SIG model identifies invariant variables for MSDA with fewer domain constraints.
problem Challenges in enforcing minimal changes across domains for MSDA.
method Subspace identification theory and variational inference.
result SIG model outperforms existing techniques on various benchmark datasets.
Work shows hallucination detection by LLMs is impossible without expert feedback.
problem Detecting hallucinations in LLMs is theoretically impossible without expert-labeled feedback.
method Investigated hallucination detection using a theoretical framework inspired by language identification.
result Automated hallucination detection is impossible for most language collections without expert-labeled feedback.
Proposes a method to identify elements in a skewness matrix for multivariate skew-elliptical distributions.
problem Label switching issue in Bayesian estimation of skewness matrix.
method Imposes a positive lower-triangular constraint and uses Bayesian sparse estimation with horseshoe prior.
result Successfully estimates the true structure of skewness dependency.
MILCCI integrates labels across categories for better understanding of multi-trial data.
problem Understanding how labels encode multi-trial observations and disentangling their effects.
method Sparse per-trial decomposition leveraging label similarities within each category.
result MILCCI identifies interpretable components and integrates label information.
Refines diagnostic prediction using causal relationships.
problem Improving accuracy of pain diagnostics prediction.
method Two approaches: 1) Inference of causal relationships, 2) Post-processing refinement.
result Potential for improving pain diagnostics prediction accuracy.
Paper proposes a new method for accurate data labeling using pairwise co-occurrences.
problem Accurate data labeling via crowdsourcing with limited data.
method Pairwise co-occurrences framework and algebraic/identifiability-enhanced algorithms.
result The approach can identify the Dawid-Skene model under realistic conditions.
This work optimizes identifying good arms in nonparametric multi-armed bandits.
problem Efficiently identifying arms with high means in nonparametric settings.
method Combining reward-maximizing sampling with a nonparametric sequential test for anytime-valid labeling.
result Achieves minimax optimal stopping times for identifying arms above a threshold.
Deep learning system speeds up wildlife species identification from camera trap images.
problem Manual review of camera trap images is slow and resource-intensive.
method Combines machine and human intelligence for active learning.
result Matches state-of-the-art accuracy with minimal manual labels.
Paper improves phase identification in power systems using information theory.
problem Improving supervised learning accuracy in phase identification.
method Developed two new techniques based on information theory.
result Significant improvement in phase identification accuracy (e.g., from 51.7% to 97.3%).
Paper proposes semi-supervised learning for EEG analysis.
problem Reducing workload and delays in analyzing large unlabeled EEG datasets.
method Semi-supervised deep learning algorithm using minimal labeled data.
result Predictions can be made with as little as 5 labeled examples.
A new anomaly detection method using partial identification.
problem Detecting anomalies in large datasets.
method Partial Identification framework and PIDScore geometric anomaly measure.
result PIDForest outperforms other methods in anomaly detection.
A simple unsupervised approach for cross-domain person re-identification.
problem Challenges in domain adaptation for person re-identification.
method Self-similarity Grouping (SSG) approach to learn pseudo identities from unlabeled samples.
result Significant improvement in mAP performance compared to state-of-the-art methods.
SSL framework identifies non-linear systems without labeled data.
problem System identification in non-linear environments without labeled data.
method Dynamics contrastive learning framework.
result SSL can identify non-linear dynamics in latent space.
Face recognition system trained with noisy labels.
problem Label noise in training deep learning classifiers.
method Review and apply recent methods to manage noisy annotations.
result Improved performance of face recognition system with noisy labels.
Martingale Doppelgänger-Eval benchmarks VLMs on candlestick evidence vs. trend extrapolation
problem Auditing whether VLMs use chart evidence or trend extrapolation
method Proving formal limitations and designing controlled mechanisms
result Identifying regression coefficients for evidence vs. trend
The paper proposes a method to identify latent factors from sampled and fired graph data.
problem Identifying latent factors from sampled and fired graph data.
method The paper presents a theoretical and practical approach to build an identifier of latent factor activations.
result The method successfully identifies latent factor activations from sampled and fired graph data.
HERA improves PLL by integrating heterogeneous loss and sparse-low-rank regularization.
problem Learning from data with partial labels.
method Combines heterogeneous loss and sparse-low-rank regularization.
result Achieves superior performance on artificial and real-world data.
Unified model identifies and locates thoracic abnormalities with limited annotations.
problem Accurate identification and localization of thoracic abnormalities require large annotated datasets, which are expensive to acquire.
method Unified model that simultaneously identifies and localizes abnormalities using limited location annotations.
result Unified model significantly outperforms baseline in classification and localization tasks.
This paper identifies and estimates the label noise transition matrix without ground truth labels.
problem Learning with noisy labels and identifying the noise transition matrix.
method Building on Kruskal's identifiability results, the paper characterizes the identifiability of the label noise transition matrix for the generic case at the instance level.
result The necessity of multiple noisy labels in identifying the noise transition matrix for the generic case at the instance level.
Method converts age labels into distributions to improve speaker age estimation.
problem Label ambiguity in age labels makes precise speaker age estimation challenging.
method Converts age labels into label distributions and uses label distribution learning.
result Our method outperforms baseline methods by reducing MAE by 10% on a real-world dataset.
Semi-supervised learning identifies radio signals from sparse data.
problem Lack of labeled data for radio emitter recognition.
method Combines unsupervised and supervised learning for feature learning and clustering.
result Semi-supervised learning can identify new radio signals efficiently.
Active learning reduces font classification data needs for historical documents.
problem Identifying fonts in historical documents for OCR accuracy.
method Active learning strategy using image features and bag-of-word representation.
result Combination of uncertainty and diversity sampling achieves 89% accuracy with 17% labeled data.
Study on identifying labels in pooled tests with noise and errors.
problem Identifying labels in pooled tests with noise and errors.
method Exact asymptotic threshold and information-theoretic framework for noisy and noisy-noise models.
result Noise can significantly increase the difficulty of the problem, even at low levels.
Automatically identifies vehicles from audio sensors without needing labeled data.
problem Vehicle recognition and classification from acoustic signals.
method Incremental reseeding of acoustic signatures using spectral embedding and clustering.
result Incremental reseeding accurately identifies individual vehicles from their acoustic signatures.
Automated labeling of intracranial arteries improves accuracy and efficiency.
problem Challenges in accurately labeling intracranial arteries due to variations and limited datasets.
method Graph Neural Network (GNN) combined with hierarchical refinement for improved accuracy.
result Achieved 97.5% node labeling accuracy on a testing set of 105 scans.
SELF filters noisy labels to improve deep learning performance.
problem Overfitting to noisy labels in deep learning.
method Self-ensemble label filtering (SELF) using running averages of predictions.
result SELF improves task performance by filtering noisy labels dynamically.
This paper solves deep learning's edge sensitivity issue by swapping important and irrelevant segments in synthetic data.
problem Edge sensitivity and high computational cost in deep learning classification models.
method Synthetic data with swapped segments to implicitly define receptive fields, preserving label information.
result The method drives networks to early convergence and appropriate solutions, improving person re-identification.
Machine learning models trained on indirect data labels can fail on real-world examples.
problem Validity issues in machine learning when target labels are indirectly defined.
method Identification of problematic datasets and models using a general procedure.
result Machine learning models trained on indirect data labels will fail on real-world examples.
PSICA identifies best treatments for patients with categorical therapies.
problem Identifying best treatments for patients with categorical therapies.
method Decision tree approach for subgroup identification in categorical treatment scenarios.
result Outputs a decision tree showing probabilities of best treatments for patient subgroups.
New method reduces text classification errors by learning writing style instead of content.
problem Deep neural networks learn superficial patterns specific to training data.
method Adversarial training to unlearn confounding features.
result Model generalizes better and learns writing style features.
CNNs identify stock market trend endpoints based on expert opinion.
problem Finding optimal entry and exit points for stock market trends.
method Three CNN submodels sequentially identify changepoints, locate them, and classify trends as upward, downward, or flat.
result CNNs can identify long-term trends based on expert opinion, offering a new approach to stock market analysis.
GGDA simplifies DA for large models, speeding up attribution by up to 50x.
problem Computational intensity of existing DA methods limits their applicability to large-scale models.
method Generalized Group Data Attribution (GGDA) framework attributing to groups of training points.
result GGDA achieves up to 50x speedups over standard DA methods while maintaining effectiveness.
A new method resolves non-identifiability in reward modeling using anchor labels.
problem Non-identifiability in reward modeling from pairwise preferences alone.
method Anchor-guided Variance-aware Reward Modeling (AVRM) framework.
result AVRM resolves non-identifiability and improves reward modeling performance.
The labeled stochastic block model is a random graph model representing networks with community structure and interactions of multiple types. In its simplest form, it consists of two communities of approximately equal size, and the edges are drawn and labeled at random with probability depending on whether their two en…
A new method embeds labels and group information for efficient multi-label classification.
problem Efficient multi-label classification with label sparsity and group structure.
method Identifies label groups, embeds labels and features in a low-dimensional space preserving sparsity and group structure.
result Our method outperforms state-of-the-art algorithms on benchmark datasets.
Deep neural nets improve KBC for scientific articles.
problem Lack of labelled data for keyphrase boundary classification.
method Multi-task learning with deep recurrent neural networks.
result Multi-task models significantly outperform previous methods.
Interactive framework identifies insights in power grid maps.
problem Manual identification of insights in power grid maps is laborious and expertise-dependent.
method Proposes an interactive framework using DenseU-Hierarchical VAE to learn and modulate representations.
result Framework outperforms baseline models in identifying and annotating insights.
Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.
problem Statistical inference challenges with missing at random labels and spatial dependence.
method Doubly robust estimator with cross-fit nuisances and jackknife spatial HAC variance correction.
result Asymptotically valid confidence intervals with improved finite-sample calibration.
Evaluating prediction models under covariate shift and selective labels
problem Model performance evaluation under distribution shift and selection bias
method Double machine learning
result Accurate estimation of target risk
Algorithm identifies and corrects training set bugs using trusted items.
problem Training set flaws affect machine learning models.
method Algorithm uses a combination of combinatorial and continuous optimization to identify and correct bugs in the training set.
result The algorithm can effectively identify and suggest changes to correct training set bugs.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
problem Noisy object re-identification in image datasets.
method Reframed Re-ID as a similarity task, using Siamese networks and Beta mixture models.
result Superior performance in noisy conditions compared to state-of-the-art methods.