Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3206399591,278 · Jun 202019922001200920182026
48 results for small labeled data

Binary classification improves with a small fraction of corrupted labels.

problem Binary classification with corrupted labels.
method Established corruption as a form of regularization and computed upper bounds on estimation error.
result Corruption is beneficial only up to a small fraction of the total sample, scaling with the square root of the sample size.

In many applications the process of generating label information is expensive and time consuming. We present a new method that combines active and semi-supervised deep learning to achieve high generalization performance from a deep convolutional neural network with as few known labels as possible. In a setting where a …

2018-03-03abs ↗pdf ↗

A scalable graph-based SSL method for large-scale data with few labels.

problem Challenges in semi-supervised learning with limited labeled data and large unlabeled data.
method Constructs a graph from a small set of high-dense vertexes to learn relationships and improve performance.
result Achieves good classification performance, especially with few labels.

Learning with noisy labels is one of the hottest problems in weakly-supervised learning. Based on memorization effects of deep neural networks, training on small-loss instances becomes very promising for handling noisy labels. This fosters the state-of-the-art approach "Co-teaching" that cross-trains two deep neural ne…

2019-01-14abs ↗pdf ↗

Unsupervised meta-learning improves learning from small labeled data.

problem Acquiring representations from unlabeled data for effective downstream learning.
method Develops an unsupervised meta-learning method that optimizes for task learning ability from unlabeled data.
result Simple task construction mechanisms, like clustering embeddings, lead to good performance on various downstream tasks.

New method pools labels from similar data items to improve learning from small samples.

problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.

Online active regression algorithms minimize label queries for efficient data regression.

problem Efficiently predict data points with minimal label queries in an online setting.
method Proposed online algorithms for active regression under ℓ_p loss, achieving (1+ε)-approximation with minimal label queries.
result Achieves (1+ε)-approximation with only ε^(-1) d log(nκ) label queries, matching offline methods in performance.

Method generates prototypes from small datasets for efficient learning.

problem Efficiently learning from small datasets with soft labels.
method Modular method for generating soft-label prototypical lines and Hierarchical Soft-Label Prototype k-Nearest Neighbor algorithm.
result High classification accuracy with significantly fewer prototypes than classes.

Self-training improves generalization by fitting to reliable pseudo-labels or gradually improving the classification plane.

problem Understanding how self-training improves generalization in high-dimensional Gaussian mixtures.
method Analyzing iterative self-training on binary Gaussian mixtures in the asymptotic limit.
result ST improves generalization by fitting to reliable pseudo-labels or gradually improving the classification plane.

This paper uses synthetic data to improve machine learning performance on small, imbalanced datasets.

problem Improving machine learning performance on small and imbalanced datasets.
method Generates synthetic data through convex combination and uses it in a semi-supervised learning framework with support vector machines.
result Synthetic data over-sampling supports the cluster assumption in semi-supervised learning, leading to outstanding results for small high-dimensional datasets and imbalanced learning problems.

Deep RL detects anomalies from few labeled examples and large unlabeled data.

problem Anomaly detection with limited labeled data and large unlabeled data.
method Deep reinforcement learning to optimize detection of labeled and unlabeled anomalies.
result Significantly outperforms state-of-the-art methods on 48 real-world datasets.

Paper tackles inherent risk scoring with choice-based data labeling and synthetic data collection.

problem Inconsistent expert judgments and lack of labeled data in inherent risk scoring.
method Choice-based data labeling and synthetic data collection.
result System achieves 89% accuracy on a test set of 52 examples.

One significant challenge to scaling entity resolution algorithms to massive datasets is understanding how performance changes after moving beyond the realm of small, manually labeled reference datasets. Unlike traditional machine learning tasks, when an entity resolution algorithm performs well on small hold-out datas…

2015-09-10abs ↗pdf ↗

In this paper, we propose a method for training neural networks when we have a large set of data with weak labels and a small amount of data with true labels. In our proposed model, we train two neural networks: a target network, the learner and a confidence network, the meta-learner. The target network is optimized to…

2017-11-30abs ↗pdf ↗

Deep semi-supervised anomaly detection improves fraud detection in financial markets.

problem Detecting fraud in high-frequency financial data with limited labeled examples.
method Evaluation of Deep Semi-Supervised Anomaly Detection (Deep SAD) on proprietary limit order book data.
result Deep SAD significantly improves fraud detection accuracy with minimal labeled data.

The paper analyzes consistency of graph-based semi-supervised learning methods for binary and multi-class classification.

problem Consistency of semi-supervised learning algorithms on graphs with noisy labels and well-clustered unlabelled data.
method The study examines graph-based probit and one-hot encoding methods for binary and multi-class classification, analyzing the consistency of optimization-based techniques.
result The analysis reveals insights into the rational function choice for optimization, improving the consistency of semi-supervised learning algorithms.

Paper tackles label noise in large datasets, purifying noisy data with a nonparametric framework.

problem Label noise in large-scale datasets with coarse labels.
method Develops a model-agnostic nonparametric framework for classification.
result Framework purifies noisy data using a small clean dataset and manages ambiguous samples.

Improves NILM with multi-label SRC, outperforming state-of-the-art.

problem Non-intrusive load monitoring (NILM) for energy disaggregation.
method Modified multi-label sparse representation based classification (SRC).
result Significant improvement over state-of-the-art techniques with minimal training data.

PSDR improves robustness against noisy labels by penalizing KL divergence between similar inputs.

problem Robust training of DNNs in datasets with noisy labels.
method Introduces PSDR, a manifold regularizer that penalizes KL divergence between similar inputs.
result Significantly improves robustness against noisy labels on benchmark datasets.

Human labeling of data can be very time-consuming and expensive, yet, in many cases it is critical for the success of the learning process. In order to minimize human labeling efforts, we propose a novel active learning solution that does not rely on existing sources of unlabeled data. It uses a small amount of labeled…

2018-05-11abs ↗pdf ↗

PPI++ uses machine learning predictions to improve inference from small datasets.

problem Efficient inference from small labeled datasets with high-quality predictions.
method Adapts prediction-powered inference (PPI) to compute confidence sets for any parameter dimensionality.
result Improves classical intervals using only labeled data, always yielding better results.

A new method classifies high-dimensional images with minimal labels using diffusion geometry.

problem Classifying high-dimensional images efficiently and accurately.
method Spatially-regularized nonlinear diffusion geometry for clustering and active learning.
result High-accuracy labelings achieved with a very small number of training labels.

GGAN improves audio representation learning with fewer labels.

problem Learning representations for specific tasks from unlabelled data.
method Guided Generative Adversarial Neural Network (GGAN).
result GGAN learns better representations with fewer labelled data.

Paper proposes a method to adapt classifiers using complementary labels instead of true labels.

problem Training classifiers with true labels from the source domain is costly and sometimes impossible.
method Proposes a novel setting with complementary labels and a complementary label adversarial network (CLARINET).
result CLARINET significantly outperforms baselines on handwritten digits and object recognition tasks.

TrustNet robustly learns noise patterns from trusted data to improve weakly-supervised classification.

problem Robustness to label noise in weakly-supervised learning.
method TrustNet learns noise patterns from trusted data, then trains a robust classifier using these patterns.
result TrustNet outperforms state-of-the-art methods in robustness to various noise patterns.

Class imbalance is an intrinsic characteristic of multi-label data. Most of the labels in multi-label data sets are associated with a small number of training examples, much smaller compared to the size of the data set. Class imbalance poses a key challenge that plagues most multi-label learning methods. Ensemble of Cl…

2018-07-30abs ↗pdf ↗

Active learning method for high-dimensional data using diffusion processes.

problem High-dimensional data labeling with limited labels.
method Learning intrinsic data geometries through diffusion processes on graphs, using diffusion distances to parametrize low-dimensional structures.
result The method achieves high-accuracy labelings with only a small number of carefully chosen labels.

Discriminative clustering learns from both labeled and unlabeled data.

problem Clustering complex datasets with limited labeled data.
method Gradient-based stochastic training and optimal transport with entropic regularization.
result The method can learn feature representations even in fully unsupervised settings.

Automatically mined rules from dependency parsing help neural models learn from less labeled data.

problem Lack of labeled data for aspect and opinion term extraction.
method Automatically mined rules from dependency parsing, applied to auxiliary data, combined with human-annotated data.
result Neural models achieve better performance than state-of-the-art with mined rules and auxiliary data.

This paper proposes a new method to select labeled data points using VAEs for active learning.

problem High cost of acquiring labels in supervised machine learning.
method Data-driven approach using Variational Autoencoders (VAEs) to select a diverse core-set in a low-dimensional latent space.
result Improvement in accuracy over related techniques, highlighting the representation power of generative modeling.

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

A semi-supervised learning method using predefined class centroids for image classification.

problem Reducing the need for labeled data in deep learning.
method Use a small number of labeled samples and data augmentation on unlabeled samples. Constrain all samples to predefined evenly-distributed class centroids (PEDCC) using loss functions.
result Achieves state-of-the-art results with minimal labeled data.