A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.
Data stream clustering tackles real-time data processing challenges.
problem Real-time processing of data streams with less prior information.
method Review of data stream clustering algorithms and their characteristics.
result Comparison and analysis of data stream clustering algorithms.
Bayesian tensor train method recovers streaming data with high accuracy.
problem Recovering high-order, incomplete, and noisy streaming data.
method Bayesian tensor train decomposition using streaming variational Bayes method.
result The proposed SPTT algorithm excels in recovering streaming data compared to state-of-the-art methods.
stream-learn is a Python library for analyzing data streams with various drift types.
problem Analyzing drifting and imbalanced data streams.
method Synthetic data stream generator, evaluation methodologies, and imbalanced binary classification metrics.
result Efficient implementation of classifiers for data stream analysis.
New method rebalances evolving data streams incrementally.
problem Incremental rebalancing of evolving data streams.
method Proposes a new streaming approach for rebalancing data streams online.
result Outperforms existing approaches in rebalancing data streams.
New features capture the order of data streams.
problem Handling ordered moments in massive data streams.
method Introducing features for ordered moments.
result Theoretical guarantees for learning algorithms.
Context improves one-class classifiers in dynamic data streams.
problem Improving one-class classification in data streams with limited training data.
method Proposes using context to guide one-class classifier learning in data streams, presenting three frameworks.
result The use of context can improve the performance of streaming one-class classifiers.
Scikit-multiflow is a Python framework for multi-output/stream data mining.
problem Handling multi-output/stream data efficiently.
method Multi-output/multi-label stream data mining framework with state-of-the-art methods.
result Enables democratization of stream learning research.
Sketches linear classifiers using Weight-Median Sketch for efficient data stream analysis.
problem Efficiently learning and analyzing data streams with limited memory.
method Introduces Weight-Median Sketch for compressed linear classifier learning over data streams.
result Memory-limited execution of various analyses over streams, including feature selection and mutual information estimation.
Dynamic Model Tree improves online learning for evolving data streams.
problem Effective and transparent machine learning on data streams is challenging.
method Revisit Model Trees for data stream applications, introducing Dynamic Model Tree.
result Dynamic Model Tree reduces the number of splits and outperforms state-of-the-art models.
New framework improves fraud prediction with incremental data balancing for massive data streams.
problem Class imbalance problem in massive imbalanced data streams.
method Incremental data balancing framework using Racing Algorithm for automated balancing and Random Forest for classification.
result Better results than Batch mode on European Credit Card dataset.
Adapts DPMM for fast streaming data clustering.
problem Clustering streaming data with time-dependent statistics.
method Adapts DPMM and sampling-based inference for online clustering.
result Obtains state-of-the-art results in speed and accuracy.
CROC identifies the earliest-changing stream as the root cause in multi-stream data.
problem Distribution-free root cause analysis in multi-stream data with unknown distributional changes.
method Conformal p-values and finite-sample valid confidence sets.
result CROC efficiently isolates the root cause under minimal assumptions.
DSSCN improves lifelong learning of non-stationary data streams through adaptive network construction.
problem Lifelong learning of non-stationary data streams with efficient and adaptive models.
method Deep stacked stochastic configuration network (DSSCN) with self-constructing deep stacked network structure and adaptive hidden unit parameters.
result DSSCN outperforms existing data stream algorithms in continual learning of non-stationary data streams.
New algorithm clusters streaming data efficiently.
problem Challenges in clustering streaming data.
method Online clustering algorithm for unknown number of clusters.
result Produces partitions close to full data clustering.
TSK-Streams learns fuzzy rules from data streams.
problem Adaptive learning from evolving data streams.
method Combines AMRules principles with fuzzy rule advantages.
result TSK-Streams performs highly competitively in experiments.
This research generates synthetic data streams for handling concept drifts and novel classes.
problem Handling concept drifts and novel classes in dynamic data streams.
method Synthetic data stream generation for both concept drifts and novel classes.
result Demonstrates the effectiveness of unsupervised drift detectors in open set recognition.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
SDF adapts Deep Forest for evolving data streams with active learning.
problem Adapting Deep Forest for evolving data streams.
method Streaming Deep Forest (SDF) with Augmented Variable Uncertainty (AVU) active learning.
result SDF with AVU outperforms other methods trained with all instances by 70% labeling budget.
Paper improves Oja's algorithm for Markovian data streams.
problem Estimating the top eigenvector of a covariance matrix from Markovian data.
method Improves Oja's algorithm for streaming PCA with Markovian dependence.
result First sharp rate for Oja's algorithm on entire data stream, removing sample size dependence.
Estimates customer segments from continuous marketing data streams.
problem Analyzing large, continuously updated marketing data streams.
method oFMLR: online estimation of finite mixture of logistic regression models.
result oFMLR provides interpretable customer segment clustering.
S-Isomap++ learns from streaming data on multiple manifolds.
problem Streaming data from multiple manifolds with irregular sampling.
method Proposes S-Isomap++ for multi-manifold learning from streaming data.
result Successfully learns from multiple intersecting manifolds in streaming data.
Bayesian model identifies outliers and determines tensor rank in streaming data.
problem Outliers and over-fitting in streaming tensor factorization.
method Variational Bayesian Inference for robust tensor rank determination and outlier identification.
result Model accurately identifies sparse outliers and determines tensor rank.
Efficiently identifies stable manifolds from streaming data.
problem Learning reliable manifolds from streaming data is computationally expensive.
method Presented error metrics and S-Isomap algorithm for efficient manifold learning.
result Identifies the transition point for stable manifold learning.
Paper proposes a method to estimate coverage in data streams.
problem Estimating coverage in data streams with limited storage.
method Modified CVM algorithm for estimating coverage in streaming settings.
result The method efficiently estimates coverage in data streams.
Melanie improves predictive performance in non-stationary data streams by transferring knowledge between multiple sources.
problem Concept drift in data streams leads to poor predictive performance.
method Melanie uses multiple sub-classifiers to learn different aspects from various sources and compose an ensemble for the target concept.
result Melanie improves predictive performance over existing algorithms by leveraging multiple sources.
A method detects changes in heterogeneous data streams over graph nodes.
problem Detecting changes in data streams from nodes of a graph.
method Online non-parametric method using likelihood-ratio estimation.
result The method accurately identifies change-points in real-world applications.
History PCA improves streaming PCA by retaining past data for better convergence.
problem Limited memory in small devices for high-dimensional data.
method History PCA algorithm that uses O(Bd) memory with B≈10 and O(d) memory with B≈1. result History PCA converges faster and performs better than existing methods.
Non-parametric method predicts multi-stream longitudinal data evolution.
problem Predicting the evolution of multi-stream longitudinal data for an in-service unit.
method Decomposes each stream into eigenfunctions and FPC scores, uses Gaussian process prior and empirical Bayesian updating.
result Framework outperforms state-of-the-art approaches and achieves high predictive accuracy.
A new method for estimating SW from streaming data.
problem Estimating Wasserstein distance from sample streams.
method Introducing a streaming estimator of the 1DW and applying it to all projections.
result Stream-SW achieves more accurate approximation of SW than random subsampling.
CURIE uses cellular automata to detect concept drift in data streams.
problem Detecting changes in data distribution (concept drift) in data streams.
method CURIE represents data stream distribution in a cellular automata grid and uses its neighborhood rule to detect changes.
result CURIE, when hybridized with base learners, performs competitively in detection metrics and classification accuracy.
New method learns manifolds from streaming data, detecting shifts.
problem Streaming data with sudden changes or gradual drifts.
method Gaussian Process Regression (GPR) with manifold-specific kernel.
result GPR can effectively learn manifold representations in streaming data.
Differentially private ensemble classifiers adapt to data streams while protecting privacy.
problem Adapting to evolving data characteristics while protecting private information.
method Unbounded ensemble updates, model agnostic approach.
result Outperforms competitors on various privacy, drift, and distribution settings.
Paper discovers process models from online event streams.
problem Discovering process models from continuous event streams.
method Generic architecture for process discovery in event streams.
result The proposed architecture enables process discovery from event streams.
Visual analytics tool detects and corrects concept drift in data streams.
problem Concept drift causes inaccurate predictions in evolving data.
method DriftVis combines drift detection and visualization.
result Visual analytics supports detection, examination, and correction of concept drift.
Detects synchronized behavior in streaming data.
problem Tracking synchronized behavior in time-stamped tuples.
method AugSplicing algorithm for streaming dense block detection.
result Effective and robust in detecting anomalous behavior.
Paper addresses challenges in benchmarking stream learning algorithms with real-world data.
problem Lack of publicly available non-stationary real-world datasets for evaluating stream algorithms.
method Proposes a new public data repository for benchmarking stream algorithms with real-world data.
result Mitigates problems related to dataset choice in experimental evaluation of stream classifiers and drift detectors.
Algorithm learns principal curves from data streams in a sequential manner.
problem Summarizing large data streams using PCA is challenging due to theoretical and algorithmic issues.
method Proposes a novel sequential algorithm for learning principal curves from data streams.
result Supports regret bounds with optimal sublinear remainder terms.
Enhash detects concept drift in data streams quickly and efficiently.
problem Detecting abrupt, gradual, virtual, or recurring events in data streams.
method Uses projection hash to insert incoming samples and detects concept drift.
result Enhash has competitive performance and moderate resource requirements compared to existing ensemble learners.
Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…
New algorithm detects and handles concept drift in data streams.
problem Handling concept drift in data streams for timely predictions.
method Hybrid Forest algorithm combining Hoeffding Trees and fast startup.
result The algorithm outperforms other methods in classification and regression tasks.
Paper combines RNN and signatures for learning functions on streamed multimodal data.
problem Learning functions on streamed multimodal data.
method Hybrid Logsig-RNN algorithm combining signatures and RNN.
result Hybrid algorithm achieves outstanding accuracy with superior efficiency and robustness.
New robustness certificates for streaming models with a sliding window.
problem Applying robustness certificates to streaming data with correlated inputs.
method Deriving robustness certificates for models using a sliding window over a sequence of potentially correlated inputs.
result Guarantees hold for the average model performance across the entire stream, independent of stream size.
This article surveys online machine learning in big data streams.
problem Limited storage for past data in data streams.
method Distributed software architectures and libraries for efficient algorithms.
result Overview of classification, regression, recommendation, and unsupervised models for streaming data.
This paper reviews methods for distributed training of machine learning models from high-rate streams.
problem Training machine learning models from high-rate distributed streams in a compute- and bandwidth-limited setting.
method Recently developed methods for large-scale distributed stochastic optimization.
result There exist regimes where systems can learn from distributed, streaming data at order-optimal rates.
A new algorithm for K-means clustering in evolving data streams.
problem Clustering of continuously arriving data in streaming scenarios with concept drift.
method Formal definition of Streaming K-means, surrogate error function, algorithm for minimizing surrogate error. result The surrogate error function effectively approximates the Streaming K-means error. Proposes methods to make data streams fair without fixing a model.
problem Fairness of data-driven models in evolving data streams.
method Modifies input data to ensure fair outcomes for any classifier.
result Improves predictive performance and low discrimination scores over time.
A framework selects the best (classifier, detector) pair for evolving data streams.
problem Selecting the best (classifier, detector) pair for data streams evolving over time.
method Reservoir of diverse adaptive learners and stacking fast Hoeffding drift detection methods.
result The best (classifier, detector) pair evolves as the stream evolves and is selected by the framework.