Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Dec 199219922001200920182026
48 results for label set size

Subset LLDA improves scalability for large label sets in multi-label classification.

problem Scalability issues in Labeled Latent Dirichlet Allocation (LLDA) for large label sets.
method Subset LLDA, a simple variant of LLDA, addressing scalability issues.
result Subset LLDA outperforms LLDA and extreme multi-label classification algorithms on large label sets.

Binary classification improves with a small fraction of corrupted labels.

problem Binary classification with corrupted labels.
method Established corruption as a form of regularization and computed upper bounds on estimation error.
result Corruption is beneficial only up to a small fraction of the total sample, scaling with the square root of the sample size.

DAL improves active learning for neural networks with large batch sizes.

problem Efficiently choosing examples to label for neural networks with large batch sizes.
method DAL treats active learning as a binary classification task to make labeled and unlabeled sets indistinguishable.
result DAL performs on par with state-of-the-art methods in medium and large query batch sizes.

Study shows resampling labels improves classifier performance in noisy data.

problem Balancing sample size vs label reliability in noisy data.
method Comparing different validation strategies and analyzing MNIST database with varying noise levels.
result Classifier performance declines with high incorrect labels, highlighting the importance of resampling.

Paper proposes a cost-sensitive conformal training method with provably controllable learning bounds.

problem Uncertainty quantification and learning bounds in conformal prediction.
method Cost-sensitive conformal training algorithm that minimizes the expected size of prediction sets using rank weighting.
result Theoretical analysis shows tightness between weighted objective and expected size of conformal prediction sets.

Study improves binary classification with multiple corrupted samples.

problem Binary classification with multiple corrupted training samples.
method Minimizes weighted combination of corruption-corrected empirical risks.
result Optimal weights are functions of sample sizes and corruption degrees.

New approach to fairness in machine learning models using conformal prediction.

problem Fairness in machine learning models' downstream decision-making.
method Theoretical derivation and empirical evaluation of label-clustered conformal prediction.
result Label-clustered conformal prediction often provides a favorable balance between utility and substantive fairness.

MAS scores cluster size consistency from points, robust to label changes.

problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.

In most classification tasks there are observations that are ambiguous and therefore difficult to correctly label. Set-valued classifiers output sets of plausible labels rather than a single label, thereby giving a more appropriate and informative treatment to the labeling of ambiguous instances. We introduce a framewo…

2016-09-02abs ↗pdf ↗

Proposes a method to improve class-conditional conformal prediction for many classes.

problem Weak guarantees for specific classes in classification problems.
method Clusters classes with similar conformal scores and performs conformal prediction at the cluster level.
result Clustered conformal typically outperforms existing methods in class-conditional coverage and set size metrics.

New method pools labels from similar data items to improve learning from small samples.

problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.

Paper proposes scalable multi-label classification for edge devices using CNN.

problem Challenges in deploying multi-label CNN models on edge devices due to high computation and memory requirements.
method Extends existing multi-label classification methods with a single CNN model and multiple loss and accuracy layers.
result Achieves comparable accuracy with 1.8x less MACC operations, 0.97x reduction in latency and 0.5x, 0.84x, 0.97x reduction in size for generated CNN models.

GAML tackles multilabel classification over graphs using message passing and attention.

problem Multilabel classification over graphs with variable-size substructures and label-substructure relations.
method GAML uses a graph neural network that models labels as auxiliary nodes and iteratively applies message passing and attention mechanisms.
result GAML significantly outperforms other methods and provides intuitive visualizations.

In the supervised learning setting termed Multiple-Instance Learning (MIL), the examples are bags of instances, and the bag label is a function of the labels of its instances. Typically, this function is the Boolean OR. The learner observes a sample of bags and the bag labels, but not the instance labels that determine…

2011-07-11abs ↗pdf ↗

New method for estimating mean in SS inference with selection bias and decaying overlap.

problem Estimating mean in SS inference with selection bias and decaying overlap.
method Double Robust Semi-Supervised (DRSS) mean estimator.
result Consistent estimation of mean with correct specification of outcome or propensity score model.

Paper proposes a method to generate instance labels from weakly supervised data.

problem Weakly supervised instance labeling in medical image analysis.
method Uses multiple instance learning (MIL) and knowledge distillation to generate instance-level predictions.
result Significantly outperforms state-of-the-art MIL methods in instance-level prediction.

This work provides a scaling rule for model EMA optimization across batch sizes.

problem Training dynamics and performance differences across batch sizes when using model EMA.
method Developed a scaling rule for model EMA optimization, demonstrating its validity across various architectures and data modalities.
result Enabled SSL methods like BYOL to train at larger batch sizes without performance degradation.

Paper proposes a method to train deep text classification models robust to label noise.

problem Training deep text classification models with noisy labels.
method Introduces a non-linear processing layer (noise model) into CNN architecture, learned jointly with CNN weights.
result The approach enables better sentence representations and robustness to extreme label noise.

The study examines label smoothing to improve confidence calibration in fine-tuned LLMs.

problem Improving confidence calibration in fine-tuned large language models (LLMs) after instruction tuning.
method Examine various open-sourced LLMs, label smoothing, and custom kernel design.
result Label smoothing is effective in maintaining confidence calibration but faces challenges in large vocabulary LLMs.

This work reveals how label noise can cause a final ascent in neural network performance curves.

problem The impact of label noise on the performance of neural networks.
method Theoretical analysis and extensive experiments on various neural network architectures.
result Label noise can lead to a final ascent in the test loss curve, improving generalization at intermediate model widths.

Active feature selection uses mutual information to choose fewer labels for better feature selection.

problem Selecting features with limited labeled data.
method Uses active feature selection with mutual information criterion, optimizing label selection for better feature quality.
result Algorithm selects features with higher mutual information using fewer labels than the data set size.

Study robustness of conformal prediction to label noise in regression and classification.

problem Robustness of conformal prediction to label noise in regression and classification.
method Characterized robustness of conformal prediction for both regression and classification problems, extending theory to control general loss functions.
result Conformal prediction and risk-controlling techniques can achieve conservative risk over clean ground truth labels with noisy labels.

In this paper we propose strategies for estimating performance of a classifier when labels cannot be obtained for the whole test set. The number of test instances which can be labeled is very small compared to the whole test data size. The goal then is to obtain a precise estimate of classifier performance using as lit…

2016-07-09abs ↗pdf ↗

The paper compares aggregated data labels in curated and random bags for machine learning models.

problem Protecting user privacy in machine learning systems with aggregated data.
method Examined curated and random bags for training machine learning models and compared their performance.
result Gradient-based learning can be performed on aggregated data without performance degradation.

CMRM improves robustness in noisy label settings without requiring privileged knowledge.

problem Learning with noisy labels without privileged knowledge.
method Conformal Margin Risk Minimization (CMRM) framework.
result CMRM consistently improves accuracy and reduces mislabeling under various noise conditions.

Estimates learnability from small data, showing accuracy with few samples.

problem Estimating how well a model class can fit a distribution of labeled data.
method Sublinear sample size estimation for learnability, extending to non-isotropic settings.
result Accurate estimation of learnability with O(d)O(\sqrt{d}) samples, even for noisy labels.

Learning from Label Proportions (LLP) is a learning setting, where the training data is provided in groups, or "bags", and only the proportion of each class in each bag is known. The task is to learn a model to predict the class labels of the individual instances. LLP has broad applications in political science, market…

2014-02-24abs ↗pdf ↗

This paper proposes an improved active learning method using classification trees.

problem Reducing the size of training sets while maintaining high accuracy in supervised learning.
method A wrapper active learning method using a classification tree to sub-sample from low-entropy regions.
result The proposed method constructs accurate classification models even with severely restricted labeled data.

This paper defines a duality for labeled graphs and factorizations, linking them to graph embeddings and Hurwitz enumeration.

problem Understanding the relationship between labeled graphs and factorizations in symmetric groups.
method Defining a mind-body duality and interpreting it in terms of Properly Embedded Graphs and Cellularly Embedded Graphs.
result Established a connection between factorizations, labeled graphs, and graph embeddings, including applications to Hurwitz enumeration.

One significant challenge to scaling entity resolution algorithms to massive datasets is understanding how performance changes after moving beyond the realm of small, manually labeled reference datasets. Unlike traditional machine learning tasks, when an entity resolution algorithm performs well on small hold-out datas…

2015-09-10abs ↗pdf ↗

New method scales Gaussian processes with derivatives using variational inference.

problem Scaling Gaussian processes with derivative information for high-dimensional problems.
method Introducing inducing directional derivatives to sparsify derivative information using variational inference.
result Achieves fully scalable Gaussian process regression with derivatives.

The ever-increasing size of modern data sets combined with the difficulty of obtaining label information has made semi-supervised learning one of the problems of significant practical importance in modern data analysis. We revisit the approach to semi-supervised learning with generative models and develop new models th…

2014-06-20abs ↗pdf ↗

Classification is an important tool with many useful applications. Among the many classification methods, Fisher's Linear Discriminant Analysis (LDA) is a traditional model-based approach which makes use of the covariance information. However, in the high-dimensional, low-sample size setting, LDA cannot be directly dep…

2015-09-17abs ↗pdf ↗

A new method for batch prediction sets in classification problems.

problem Constructing reliable prediction sets for multiple unlabeled examples.
method Proposes a uniformly more powerful approach to batch prediction sets using specific combinations of conformal p-values.
result The proposed method provides narrower prediction sets compared to the Bonferroni correction.

A new method for efficient label retrieval in large output spaces.

problem Efficiently retrieving relevant labels for inputs with large output spaces.
method Developed a technique called Stochastic Negative Mining to address the problem of set-valued classifiers in large output spaces.
result Stochastic Negative Mining outperforms existing negative sampling approaches in experiments.