FLAME auto-labels mobile data efficiently on diverse processors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ASEs use surrogate estimation to efficiently evaluate model performance with minimal labels.
Efficient algorithms learn from coarse labels instead of fine grained ones.
ELSA efficiently adapts to label shift without post-prediction calibrations.
We study active learning of homogeneous -sparse halfspaces in under the setting where the unlabeled data distribution is isotropic log-concave and each label is flipped with probability at most for a parameter , known as the bounded noise. Even in the presence of mild la…
New algorithms reduce label collection for online prediction with expert advice.
Efficiently tests two distributions with few label queries.
Active testing reduces label costs for efficient model evaluation.
Efficiently learns from partial labels using variational inference.
Adaptive sampling detects local concept drift with limited labels.
In this work, we propose a novel framework for the labeling of entity alignments in knowledge graph datasets. Different strategies to select informative instances for the human labeler build the core of our framework. We illustrate how the labeling of entity alignments is different from assigning class labels to single…
It has been a long-standing problem to efficiently learn a halfspace using as few labels as possible in the presence of noise. In this work, we propose an efficient Perceptron-based algorithm for actively learning homogeneous halfspaces under the uniform distribution over the unit sphere. Under the bounded noise condit…
Efficiently learns halfspaces with malicious noise, near-optimal label complexity.
We study the problem of efficient PAC active learning of homogeneous linear classifiers (halfspaces) in , where the goal is to learn a halfspace with low error using as few label queries as possible. Under the extra assumption that there is a -sparse halfspace that performs well on the data ()…
An efficient algorithm identifies labels from sparse pooled data.
Being able to model correlations between labels is considered crucial in multi-label classification. Rule-based models enable to expose such dependencies, e.g., implications, subsumptions, or exclusions, in an interpretable and human-comprehensible manner. Albeit the number of possible label combinations increases expo…
Study efficient learning of robust halfspaces with noise.
We propose a framework that learns a representation transferable across different domains and tasks in a label efficient manner. Our approach battles domain shift with a domain adversarial loss, and generalizes the embedding to novel task using a metric learning-based approach. Our model is simultaneously optimized on …
Extreme Multi-label classification (XML) is an important yet challenging machine learning task, that assigns to each instance its most relevant candidate labels from an extremely large label collection, where the numbers of labels, features and instances could be thousands or millions. XML is more and more on demand in…
ETM models improve efficiency in semi-supervised logistic regression.
This work analyzes label embedding for large multiclass classification problems.
A new learning scheme improves model efficiency and performance.
This paper tackles label-efficient evaluation in extreme class imbalance.
Efficient active learning with abstention reduces label complexity exponentially.
This paper addresses image classification through learning a compact and discriminative dictionary efficiently. Given a structured dictionary with each atom (columns in the dictionary matrix) related to some label, we propose cross-label suppression constraint to enlarge the difference among representations for differe…
PDO optimizes LLM prompts without labels, improving performance.
MEC improves efficiency and robustness in semi-supervised inference.
Recent advances in machine learning have led to increased deployment of black-box classifiers across a wide variety of applications. In many such situations there is a critical need to both reliably assess the performance of these pre-trained models and to perform this assessment in a label-efficient manner (given that…
This work efficiently learns linear threshold functions from label proportions using Gaussian distributions.
Efficient method for generating adversarial examples with limited query budget.
In this dissertation, we focus on several important problems in structured prediction. In structured prediction, the label has a rich intrinsic substructure, and the loss varies with respect to the predicted label and the true label pair. Structured SVM is an extension of binary SVM to adapt to such structured tasks. I…
Active inference framework improves -statistic estimation efficiency.
Graph-based methods have been demonstrated as one of the most effective approaches for semi-supervised learning, as they can exploit the connectivity patterns between labeled and unlabeled data samples to improve learning performance. However, existing graph-based methods either are limited in their ability to jointly …
The paper explores learning from label proportions, showing differences in efficiency between LLP and PAC learning.
Efficient algorithm for halfspaces with specific noise conditions.
Efficiently selects nearest neighbors for labeling to speed up active learning.
Proposes efficient calibration for indoor localization models.
The paper explains how data augmentation improves semi-supervised learning efficiency.
Extreme multi-label text classification (XMTC) addresses the problem of tagging each text with the most relevant labels from an extreme-scale label set. Traditional methods use bag-of-words (BOW) representations without context information as their features. The state-ot-the-art deep learning-based method, AttentionXML…
In multi-label learning, each sample is associated with several labels. Existing works indicate that exploring correlations between labels improve the prediction performance. However, embedding the label correlations into the training process significantly increases the problem size. Moreover, the mapping of the label …
This paper presents privileged multi-label learning (PrML) to explore and exploit the relationship between labels in multi-label learning problems. We suggest that for each individual label, it cannot only be implicitly connected with other labels via the low-rank constraint over label predictors, but also its performa…
LACD uses unlabeled data to improve conditional diffusion models.
Crowdsourcing has become very popular among the machine learning community as a way to obtain labels that allow a ground truth to be estimated for a given dataset. In most of the approaches that use crowdsourced labels, annotators are asked to provide, for each presented instance, a single class label. Such a request c…
Efficient AL method improves CNN performance with minimal labeled data.
Study efficient active learning for halfspaces with Tsybakov noise using non-convex optimization.
Paper connects sampling and labeling biases in large-output spaces.
While deep learning has been incredibly successful in modeling tasks with large, carefully curated labeled datasets, its application to problems with limited labeled data remains a challenge. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through a combin…
This paper investigates the problem of active learning for binary label prediction on a graph. We introduce a simple and label-efficient algorithm called S2 for this task. At each step, S2 selects the vertex to be labeled based on the structure of the graph and all previously gathered labels. Specifically, S2 queries f…