Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4008001,2001,600 · Jun 202019922001200920182026
48 results for streaming label learning

A new method for online multi-label stream classification.

problem Challenges in classifying continuous data streams with concept drift and delayed labels.
method Online unsupervised incremental method based on self-organizing maps.
result The method is highly competitive in both stationary and concept drift scenarios.

New method tackles dynamic data labeling issues with limited labels.

problem Dynamic data labeling with scarce labeled instances.
method Instance exploitation technique for aggressive model adaptation.
result Aggressive model adaptation leads to better performance than standard methods.

QActor optimizes learning from noisy labeled data streams by querying experts for clean labels.

problem Learning from noisy labeled data in continuous streams with limited oracle queries.
method Combines quality models for filtering and oracle queries for true labels, dynamically adjusting query limits.
result QActor nearly matches optimal accuracy with up to 6% additional ground truth data from experts.

Confidence intervals improve decision tree accuracy in streaming data.

problem Improving decision tree accuracy in streaming data with confidence intervals.
method Deriving accurate confidence intervals for decision tree splitting criteria and extending to selective sampling.
result Confidence intervals enhance decision tree accuracy and reduce labeling costs.

ATL learns from many streaming processes without labeled data.

problem Knowledge transfer across many streaming processes with covariate shift and drifts.
method Autonomous transfer learning with generative and discriminative phases, KL divergence optimization, and elastic network structure.
result Improved performance and faster training speed compared to existing methods.

ParsNet tackles weakly supervised data streams with a self-evolving deep neural network.

problem Weakly supervised data streams hinder existing data stream algorithms.
method ParsNet uses a self-labelling strategy with hedge (SLASH) and a closed-loop configuration of generative and discriminative training processes.
result ParsNet outperforms other methods in high-dimensional data streams and infinite delay simulations.

An active learner chooses which data points to label to minimize cost and error in streaming data.

problem Efficiently labeling streaming data points with limited labeling costs.
method Formalizes the problem with a loss function, designs an algorithm with a time and cost dependent threshold, and provides upper and lower bounds.
result The algorithm achieves a worst-case upper bound of O~(B13K13T23)\widetilde{O}(B^{\frac{1}{3}} K^{\frac{1}{3}} T^{\frac{2}{3}}) on the loss after TT rounds.

Memory augmented neural networks improve active learning for one-shot predictions.

problem Scarcity and cost of labeled training data in deep architectures.
method Memory augmented neural networks and Class Margin Sampling (CMS) for reinforcement learning.
result The proposed method outperforms existing baselines in label predictions and reduces label requests.

Paper proposes a semi-supervised method for detecting concept drift in streaming environments.

problem Detecting concept drift in streaming environments with limited labeled data.
method Utilizes density estimation of posterior probabilities in partially labeled streaming data.
result Demonstrates superior concept drift detection in streaming environments with limited labeled data.

New algorithm robust to label corruptions in active learning.

problem Active learning under unknown adversarial label corruptions.
method Proposed a new active learning algorithm that is provably correct without assumptions on corruptions.
result Achieves minimax label complexity in non-corrupted setting and only requires additional labels to achieve desired accuracy in corrupted setting.

New algorithm reduces label queries in online learning with bounded errors.

problem Minimizing label queries while limiting prediction errors in streaming data.
method Disagreement-based online learning algorithm for a general hypothesis space under Tsybakov noise.
result The proposed algorithm achieves an optimal label complexity of O(dT22α2αlog2T)O(dT^{\frac{2-2α}{2-α}}\log^2 T) with a matching lower bound.

New method combines deep learning and streaming learning for better incremental learning.

problem Catastrophic forgetting in deep neural networks when updated incrementally.
method Combining streaming linear discriminant analysis with deep learning.
result Outperforms incremental batch learning and streaming learning on ImageNet and CORe50.

Paper classifies multiple video sources in encrypted tunnels using NLP-inspired features.

problem Traffic classification in encrypted video streams.
method Deep learning with a novel NLP-inspired feature for multi-label classification.
result The method achieves high performance on binary and multilabel classification tasks.

Novel semi-supervised method for online structure learning in noisy data streams.

problem Discovering complex relations in noisy data streams with limited labelled data.
method Combines graph-cut minimization and first-order logic for online, single-pass label completion.
result Improves accuracy of structure learning system by completing missing labels.

Paper explores active learning strategies for real-time credit card fraud detection.

problem Challenges in labeling and imbalanced transaction data for real-time fraud detection.
method Investigates active learning strategies for querying unlabeled transactions, comparing supervised, semi-supervised, and unsupervised approaches.
result Highlights an exploitation/exploration trade-off for active learning in fraud detection.

We consider the problem of learning convex aggregation of models, that is as good as the best convex aggregation, for the binary classification problem. Working in the stream based active learning setting, where the active learner has to make a decision on-the-fly, if it wants to query for the label of the point curren…

2015-03-28abs ↗pdf ↗

LdSM builds efficient multi-label decision trees with logarithmic depth.

problem Efficiently annotate data points with relevant subsets of labels from a large label set.
method Develops LdSM algorithm for multi-label decision trees with logarithmic depth, optimizing a novel objective function for balanced splits and high class purity.
result Minimizing the proposed objective function leads to pure and balanced data splits, achieving high prediction accuracy and low prediction time.

This work tackles continual learning with semi-supervised data, showing that even with minimal labeled data, performance can match full-supervised methods.

problem Training deep networks on a stream of tasks without forgetting, especially when labeled data is scarce.
method Designing a novel CSSL method that leverages metric learning and consistency regularization to learn from both labeled and unlabeled data.
result Our method outperforms state-of-the-art methods trained with full supervision, achieving comparable performance with only 25% labeled data.

DBULL learns new clusters without forgetting past knowledge in streaming unlabelled data.

problem Challenges in Unsupervised Lifelong Learning with evolving data distributions and class labels.
method Bayesian framework for incremental learning, Deep Bayesian Unsupervised Lifelong Learning (DBULL) algorithm, knowledge preservation mechanism, automatic cluster discovery.
result DBULL can progressively discover new clusters without forgetting past knowledge in unlabelled data.

Research explores unsupervised methods for detecting vessel behavior changes in real-time data streams.

problem Detecting shifts in vessel behavior for maritime traffic monitoring.
method Investigates unsupervised and semi-supervised change detection methods.
result Identifies shifts in vessel behavior for unusual events detection.

Meta-learning framework for few-shot one-class classification using order-equivariant networks.

problem Few labeled examples for positive class in one-class classification tasks.
method Order-equivariant networks for meta-learning a binary classifier conditioned on positive examples.
result Meta-learning framework outperforms baselines on unseen synthetic streams.

Paper improves anomaly detection using tree-based ensembles with active learning.

problem Configuring anomaly detectors with true labels to minimize false positives.
method Develops batch and streaming active learning algorithms for tree-based ensembles.
result Significantly more anomalies discovered with active learning compared to baselines.

Approach to detect and adapt to concept drift in unlabeled streaming data.

problem Detect and adapt to concept drift in high-dimensional, noisy, low-context data.
method Density-based clustering for virtual drift and weak supervision for real drift.
result 90% precision in detecting and adapting to concept drift for 4 years after initial deployment.