CROC identifies the earliest-changing stream as the root cause in multi-stream data.
problem Distribution-free root cause analysis in multi-stream data with unknown distributional changes.
method Conformal p-values and finite-sample valid confidence sets.
result CROC efficiently isolates the root cause under minimal assumptions.
This paper reviews methods for distributed training of machine learning models from high-rate streams.
problem Training machine learning models from high-rate distributed streams in a compute- and bandwidth-limited setting.
method Recently developed methods for large-scale distributed stochastic optimization.
result There exist regimes where systems can learn from distributed, streaming data at order-optimal rates.
A new method for estimating SW from streaming data.
problem Estimating Wasserstein distance from sample streams.
method Introducing a streaming estimator of the 1DW and applying it to all projections.
result Stream-SW achieves more accurate approximation of SW than random subsampling.
Paper develops streaming, distributed inference for BNP models.
problem Handling the lack of inherent ordering in BNP models.
method Combinatorial optimization over component correspondences for asynchronous updates.
result Efficient solution to correspondence problem enables practical scalability.
Method learns Bayesian networks from distributed streaming data with reduced communication.
problem Learning and maintaining machine learning models over distributed, streaming data.
method Communication-efficient method for continuously learning Bayesian networks over a distributed stream.
result Exponential reduction in communication compared to baseline approaches.
Develops new Bayesian approach for analyzing streaming data.
problem Bayesian inference on streaming data is challenging.
method Population variational Bayes, approximating population posterior.
result Approximates population posterior for streaming data.
Visual analytics tool detects and corrects concept drift in data streams.
problem Concept drift causes inaccurate predictions in evolving data.
method DriftVis combines drift detection and visualization.
result Visual analytics supports detection, examination, and correction of concept drift.
This article surveys online machine learning in big data streams.
problem Limited storage for past data in data streams.
method Distributed software architectures and libraries for efficient algorithms.
result Overview of classification, regression, recommendation, and unsupervised models for streaming data.
New method learns manifolds from streaming data, detecting shifts.
problem Streaming data with sudden changes or gradual drifts.
method Gaussian Process Regression (GPR) with manifold-specific kernel.
result GPR can effectively learn manifold representations in streaming data.
Paper proposes a distributed method to estimate principal eigenvector from high-rate streaming data.
problem Estimating principal eigenvector from high streaming data rate.
method Distributed Krasulina (D-Krasulina) and mini-batch extension (DM-Krasulina) methods.
result Achieves optimal estimation error rates under high streaming conditions.
Sentinel analyzes Twitter streams in real-time with high accuracy.
problem Real-time analytics on fast data streams with minimal delay and high accuracy.
method Distributed system Sentinel using Apache Storm and SpaceSaving for summary storage.
result Sentinel achieves high analytical accuracy on Twitter streams.
CURIE uses cellular automata to detect concept drift in data streams.
problem Detecting changes in data distribution (concept drift) in data streams.
method CURIE represents data stream distribution in a cellular automata grid and uses its neighborhood rule to detect changes.
result CURIE, when hybridized with base learners, performs competitively in detection metrics and classification accuracy.
LOBDIF predicts limit order book events using a diffusion model.
problem Predicting the timing and type of events in a dynamic market system.
method LOBDIF uses a diffusion model to learn the complex time-event distribution in limit order book streams.
result LOBDIF significantly outperforms existing methods in real-world data experiments.
New algorithms for fair data summarization in massive data models.
problem Fair k-center problem in massive data. method Streaming and distributed algorithms with provable guarantees.
result First distributed algorithm with constant approximation ratio.
A new framework detects concept drift in streaming data.
problem Detecting distributional changes in non-stationary data streams.
method Treating model parameters as random variables, ERICS uses information theory measures to identify concept drift.
result ERICS effectively detects concept drift compared to existing methods.
A method detects changes in heterogeneous data streams over graph nodes.
problem Detecting changes in data streams from nodes of a graph.
method Online non-parametric method using likelihood-ratio estimation.
result The method accurately identifies change-points in real-world applications.
Paper improves Oja's algorithm for Markovian data streams.
problem Estimating the top eigenvector of a covariance matrix from Markovian data.
method Improves Oja's algorithm for streaming PCA with Markovian dependence.
result First sharp rate for Oja's algorithm on entire data stream, removing sample size dependence.
This research sets limits on robust learning from data with errors.
problem Learning from data with malicious errors and outliers.
method Lower bounds on communication and space complexities for robust learning.
result Gaining robustness in learning usually increases complexity.
A method for learning rankings in non-stationary data streams.
problem Learning preferences in a population that changes over time.
method Generalized Borda algorithm for non-stationary ranking streams.
result Bounds on the minimum number of samples required to output the ground truth.
StreamEnsemble dynamically selects ML models for ST data streams to improve predictive accuracy.
problem Predictive queries over spatiotemporal data streams are challenging due to varying distributions and patterns.
method Dynamic selection and allocation of ML models based on time series distributions and characteristics.
result Significantly outperforms traditional ensemble and single model approaches, reducing prediction error by over 10 times.
Paper detects changes in graph-based data streams using likelihood-ratios.
problem Detecting changes in synchronized graph-based data streams.
method Kernel-based likelihood-ratio estimation over graph nodes.
result Effective detection and localization of change-points.
Develops an online learning method for large/streaming data.
problem Prediction in large/streaming data sets.
method Covariance-fitting methodology for online learning.
result Predictor with desirable properties: linear runtime, constant memory, no local minima, prunes redundant dimensions.
Paper proposes a beta distribution method for detecting concept drift in adaptive classifiers.
problem Adaptive classifiers need to detect and respond to concept drift in real-time data streams.
method The paper introduces a beta distribution model to monitor model error and identify abnormal behavior as drift.
result The method effectively detects abrupt changes in model error, improving classifier performance.
CDRE estimates density ratios in streaming data without historical samples.
problem Online learning with shifting data distributions.
method Iterative estimation of density ratios between initial and current distributions.
result CDRE outperforms standard DRE in estimating divergences between distributions.
New online method estimates OT distances from sample streams.
problem Computing OT distances between arbitrary distributions.
method Online Sinkhorn algorithm using iterative enrichment of non-parametric representation.
result Consistent estimation of true regularized OT distance with nearly-O(1/n) sample complexity.
Bayesian nonparametric CMS improves frequency estimation for power-law data.
problem Estimating frequencies of low-frequency tokens in power-law data streams.
method Developed a learning-augmented count-min sketch using a normalized inverse Gaussian process prior.
result The approach achieves remarkable performance in estimating low-frequency tokens.
DDG-DA predicts future data distribution to adapt models for predictable concept drift.
problem Adapting models to streaming data with predictable concept drift.
method Train a predictor to forecast future data distribution, generate training samples, and train models on them.
result Significant improvement on multiple models in real-world tasks.
New findings show concept drift in data streams implies temporal dependence, affecting model design.
problem Concept drift in data streams challenges existing model design and deployment approaches.
method Developed gradient-descent methods for continuous adaptation without explicit drift detection.
result Gradient-descent methods offer major advantages in accuracy and efficiency for concept-drifting streams.
Bayesian model updates data streams with hierarchical priors.
problem Continuous model updating and adapt to changes in data distribution.
method Non-conjugate hierarchical priors and variational inference.
result Validated on real data sets, demonstrating adaptability.
Low-precision streaming PCA estimates the leading eigenvector with limited precision.
problem Estimating the leading eigenvector in a streaming setting with limited precision.
method Oja's algorithm with linear and nonlinear stochastic quantization.
result A batched version of the quantized variants achieves the lower bound on quantization error up to logarithmic factors.
Efficiently tests discrete distributions with limited memory and communication.
problem Testing discrete distributions with constraints on memory and communication.
method Developed efficient algorithms for uniformity/identity and closeness testing in streaming and distributed models.
result Nearly-tight lower bounds on sample complexity and communication cost for uniformity testing.
This paper reviews techniques for distributed learning with non-convex models.
problem Optimizing non-convex models in distributed learning.
method Selective review of recent techniques for batch and streaming data.
result Trade-offs between computation and communication costs in distributed learning.
Paper addresses challenges in benchmarking stream learning algorithms with real-world data.
problem Lack of publicly available non-stationary real-world datasets for evaluating stream algorithms.
method Proposes a new public data repository for benchmarking stream algorithms with real-world data.
result Mitigates problems related to dataset choice in experimental evaluation of stream classifiers and drift detectors.
A scalable algorithm for computing Wasserstein barycenters of streaming data.
problem Aggregating data from different, possibly non-identically distributed sources.
method Parallel, semi-discrete algorithm for continuous input distributions.
result Robust, streaming Wasserstein barycenter estimate that tracks nonstationary distributions.
Paper analyzes Async-MSGD for nonconvex problems using streaming PCA.
problem Understanding convergence properties of Async-MSGD for nonconvex problems.
method Diffusion approximation to analyze Async-MSGD for streaming PCA.
result Async-MSGD requires reduced momentum for acceleration through asynchrony.
New algorithms for interactive learning in pools vs streams.
problem Comparing pool-based and stream-based interactive learning.
method Developed algorithms for both settings and analyzed their performance.
result A maximal gap exists between pool and stream settings in active learning.
DiwE uses regional distribution changes to create diverse ensemble classifiers for concept drift.
problem Handling concept drift in evolving data streams.
method DiwE measures diversity based on regional distribution disagreement and uses it to weight instances and select classifiers.
result DiwE outperforms other algorithms on various synthetic and real-world data stream benchmarks.
Differentially private ensemble classifiers adapt to data streams while protecting privacy.
problem Adapting to evolving data characteristics while protecting private information.
method Unbounded ensemble updates, model agnostic approach.
result Outperforms competitors on various privacy, drift, and distribution settings.
A new control chart detects shifts in binary data streams quickly and reliably.
problem Early detection of small shifts in multiple binary data streams.
method Cumulative Standardized Binomial EWMA (CSB-EWMA) chart with exact variance derivation.
result Adaptive control limits ensure robust detection across different data distributions.
Develops new techniques for learning from sequential data groups.
problem Learning from groups of inputs rather than individual inputs.
method Introduces feature-based and kernel-based learning techniques for sequential data.
result Achieves state-of-the-art performance on various real-world examples.
New robustness certificates for streaming models with a sliding window.
problem Applying robustness certificates to streaming data with correlated inputs.
method Deriving robustness certificates for models using a sliding window over a sequence of potentially correlated inputs.
result Guarantees hold for the average model performance across the entire stream, independent of stream size.
Dynamic Model Tree improves online learning for evolving data streams.
problem Effective and transparent machine learning on data streams is challenging.
method Revisit Model Trees for data stream applications, introducing Dynamic Model Tree.
result Dynamic Model Tree reduces the number of splits and outperforms state-of-the-art models.
Classifiers trained on data sets possessing an imbalanced class distribution are known to exhibit poor generalisation performance. This is known as the imbalanced learning problem. The problem becomes particularly acute when we consider incremental classifiers operating on imbalanced data streams, especially when the l…
A new streaming Gibbs sampling method improves online LDA model perplexity.
problem Online learning of LDA models with poor performance.
method Streaming Gibbs Sampling (SGS) as an online extension of collapsed Gibbs sampling (CGS).
result SGS achieves similar perplexity to CGS, better than SVB.
A new framework predicts hidden Markov model regimes online.
problem Efficiently identify hidden Markov model regimes in streaming data.
method Develops a predictive-first optimisation framework for streaming HMMs, approximating the full posterior predictive distribution.
result The method provides competitive prequential performance compared to Online EM and Sequential Monte Carlo.
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
problem Online monitoring of multivariate data streams for detecting changes.
method Combines Kernel-QuantTree histogram and EWMA statistic for non-parametric monitoring.
result Controls Average Run Length (ARL0) while achieving comparable detection delays.
New analysis for sampling from non-convex distributions with dependent data.
problem Sampling from non-logconcave distributions in stochastic optimization.
method Stochastic Gradient Langevin Dynamics (SGLD) with dependent data streams.
result Sharper and uniform convergence estimates in L1-Wasserstein distance. DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.