DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This article surveys online machine learning in big data streams.
problem Limited storage for past data in data streams.
method Distributed software architectures and libraries for efficient algorithms.
result Overview of classification, regression, recommendation, and unsupervised models for streaming data.
LOFS library aids in online streaming feature selection.
problem Sequentially adding dimensions in high-dimensional data.
method State-of-the-art algorithms for online streaming feature selection.
result First open-source library for online streaming feature selection.
Efficiently updates KRR for big streams with minimal redundant computation.
problem Redundant computation in incremental KRR for big data streams.
method Supports incremental/decremental processing for single and multiple samples, dividing data into batches.
result Significantly reduced computational time without sacrificing accuracy.
New algorithms minimize regret in streaming MAB with memory constraints.
problem Minimizing regret in single-pass streaming MAB with limited memory.
method Developed two algorithms with tight regret bounds for different memory sizes.
result Established tight gap-dependent regret bounds for streaming MAB.
Sentinel analyzes Twitter streams in real-time with high accuracy.
problem Real-time analytics on fast data streams with minimal delay and high accuracy.
method Distributed system Sentinel using Apache Storm and SpaceSaving for summary storage.
result Sentinel achieves high analytical accuracy on Twitter streams.
New algorithm clusters streaming data efficiently.
problem Challenges in clustering streaming data.
method Online clustering algorithm for unknown number of clusters.
result Produces partitions close to full data clustering.
Efficiently trains SVM models on large datasets using coreset technology.
problem Training large-scale SVM models efficiently on Big Data.
method Developed an algorithm to create a coreset, a small representative subset of data.
result Proved the size of coreset required for SVM models and showed its applicability to streaming data.
New model for music streaming recommends songs based on past play history.
problem Nonstationary stochastic bandit model with delay-dependent rewards.
method Ranking policies approximating optimal policy with bounded regret.
result Algorithm with O ~ ( k T ) \widetilde{\mathcal{O}}\big(\!\sqrt{kT}\big) O ( k T ) regret and O ( k ln ln T ) \mathcal{O}\big(k\ln\ln T\big) O ( k ln ln T ) switches. FROSH and DFROSH speed up online sketching hashing for big data.
problem Efficiency and scalability of hashing methods for streaming data.
method Online sketching hashing (OSH) and FasteR Online Sketching Hashing (FROSH) algorithm.
result FROSH reduces training time and maintains sketching precision.
New algorithm detects and handles concept drift in data streams.
problem Handling concept drift in data streams for timely predictions.
method Hybrid Forest algorithm combining Hoeffding Trees and fast startup.
result The algorithm outperforms other methods in classification and regression tasks.
This paper surveys smart buildings using machine learning and big data.
problem Improving comfort and efficiency in smart buildings through data analysis.
method Survey of machine learning and big data techniques for smart buildings.
result Machine learning and big data are crucial for smart building services.
This study designs a financial risk control platform using big data and machine learning.
problem Traditional risk management models are inadequate for modern financial complexities.
method Big data mining, real-time streaming data processing, statistical analysis, and precise customer behavior mining.
result The platform effectively identifies and responds to potential risks in real-time.
Extracting latent low-dimensional structure from high-dimensional data is of paramount importance in timely inference tasks encountered with `Big Data' analytics. However, increasingly noisy, heterogeneous, and incomplete datasets as well as the need for {\em real-time} processing of streaming data pose major challenge…
A new platform predicts object movements without needing specific context.
problem Predicting the next position of movable objects.
method λ-Architecture for batch and stream analytics, combining for improved accuracy.
result Combining λ-Architecture parts improves overall accuracy and performance.
Survey of algorithms for PCA and subspace tracking with missing data.
problem Handling missing data in streaming Principal Component Analysis and subspace tracking.
method Review of classical and recent algorithms with low computational and memory complexities.
result Algorithms need careful adjustment for missing data.
Survey of big data in cyber-physical systems, including data security and green challenges.
problem Managing vast amounts of data in cyber-physical systems.
method Taxonomy and overview of data collection, storage, access, processing, and analysis.
result First panoramic survey on big data for CPS, addressing cybersecurity and green challenges.
Confidence intervals improve decision tree accuracy in streaming data.
problem Improving decision tree accuracy in streaming data with confidence intervals.
method Deriving accurate confidence intervals for decision tree splitting criteria and extending to selective sampling.
result Confidence intervals enhance decision tree accuracy and reduce labeling costs.
New algorithm reduces memory usage for multi-pass bandit problems.
problem Memory-efficient multi-pass bandit algorithms for large action spaces.
method Develops a B B B -pass algorithm with O ( 1 ) O(1) O ( 1 ) memory that achieves optimal regret. result Sharp memory-regret trade-off: O ( 1 ) O(1) O ( 1 ) memory suffices for Θ ( T 1 / 2 ) Θ(T^{1/2}) Θ ( T 1/2 ) regret in B B B passes. Paper proposes an integrated M&D approach for large multistream data.
problem Inability to progress in monitoring and diagnostics due to high-dimensionality and volume of multistream data.
method Adaptive Principal Component monitoring (APC) and Principal Component Signal Recovery (PCSR).
result The integrated M&D approach enables early detection and streamlined SPC.
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
problem Designing efficient unlearning algorithms for machine learning models.
method Formalizes the problem and gives efficient unlearning algorithms for linear and prefix-sum query classes.
result Improved guarantees for stochastic convex optimization with reduced unlearning query complexity.
A fast algorithm for generalized matrix regression improves machine learning performance.
problem Efficiently solving generalized matrix regression problems in machine learning.
method Utilizes sketching technique to achieve ( 1 + ε ) (1+ε) ( 1 + ε ) relative error with sketching sizes of order $\cO(ε^{-1/2})$ . result The Fast GMR algorithm achieves better performance in symmetric positive definite matrix approximation and single pass singular value decomposition.
ESRF reduces ARF ensemble size without sacrificing accuracy.
problem Over-provisioning of ARF ensemble leads to high CPU and memory consumption.
method ESRF uses a swap and elastic component to dynamically adjust the number of classifiers.
result ESRF reduces the number of classifiers by up to one third without sacrificing accuracy.
New algorithm tightens sensitivity bounds for coresets, reducing size by up to 10x.
problem Finding efficient coresets for least-mean-squares problems.
method Algorithm based on iterative sensitivity computation and reduction to higher-dimensional subspaces.
result Significantly smaller coresets with data-dependent sensitivity bounds.
Fast online algorithm for nonparametric correlations.
problem Computing nonparametric correlations on streaming data.
method Novel online algorithm with O(1) time and memory complexity.
result 10 to 1,000 times faster than batch algorithms.
In response to the need for learning tools tuned to big data analytics, the present paper introduces a framework for efficient clustering of huge sets of (possibly high-dimensional) data. Building on random sampling and consensus (RANSAC) ideas pursued earlier in a different (computer vision) context for robust regress…
A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.
Data stream clustering tackles real-time data processing challenges.
problem Real-time processing of data streams with less prior information.
method Review of data stream clustering algorithms and their characteristics.
result Comparison and analysis of data stream clustering algorithms.
Paper uses surprisal to dynamically allocate computation between fast and slow models.
problem Dynamic allocation of computation in neural networks.
method Surprisal-based dynamic model selection.
result Model can match baseline performance with 15% fewer FLOPs.
Developed a Swiss real estate portal using machine learning and public data.
problem Creating a real estate portal without domain expertise and making it accessible.
method Continuous web crawling of real estate ads, using machine learning for price estimation.
result Random Forest algorithm provides accurate rental price estimates with a median absolute relative error of 6.57 percent.
Bayesian tensor train method recovers streaming data with high accuracy.
problem Recovering high-order, incomplete, and noisy streaming data.
method Bayesian tensor train decomposition using streaming variational Bayes method.
result The proposed SPTT algorithm excels in recovering streaming data compared to state-of-the-art methods.
New algorithm solves constrained ℓ_p regression problems efficiently.
problem Minimizing ℓ_p regression over unit vectors with constraints.
method Uses core-sets and provable constant factor approximation.
result First provable constant factor approximation algorithm.
stream-learn is a Python library for analyzing data streams with various drift types.
problem Analyzing drifting and imbalanced data streams.
method Synthetic data stream generator, evaluation methodologies, and imbalanced binary classification metrics.
result Efficient implementation of classifiers for data stream analysis.
Paper proposes a distributed algorithm for multi-label feature selection.
problem Maximizing diversity and quality in non-redundant feature selection.
method Greedy algorithm for distributed optimization of submodular plus diversity functions.
result Achieves constant factor approximation of optimal solution in big data settings.
We analyze a compression scheme for large data sets that randomly keeps a small percentage of the components of each data sample. The benefit is that the output is a sparse matrix and therefore subsequent processing, such as PCA or K-means, is significantly faster, especially in a distributed-data setting. Furthermore,…
New method rebalances evolving data streams incrementally.
problem Incremental rebalancing of evolving data streams.
method Proposes a new streaming approach for rebalancing data streams online.
result Outperforms existing approaches in rebalancing data streams.
New features capture the order of data streams.
problem Handling ordered moments in massive data streams.
method Introducing features for ordered moments.
result Theoretical guarantees for learning algorithms.
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
This paper categorizes and analyzes mobile big data from wireless networks.
problem Understanding social characteristics of mobile big data.
method Categorization and analysis of real wireless cellular network data.
result Highlight several research directions in social computing for mobile big data.
Context improves one-class classifiers in dynamic data streams.
problem Improving one-class classification in data streams with limited training data.
method Proposes using context to guide one-class classifier learning in data streams, presenting three frameworks.
result The use of context can improve the performance of streaming one-class classifiers.
Scikit-multiflow is a Python framework for multi-output/stream data mining.
problem Handling multi-output/stream data efficiently.
method Multi-output/multi-label stream data mining framework with state-of-the-art methods.
result Enables democratization of stream learning research.
Sketches linear classifiers using Weight-Median Sketch for efficient data stream analysis.
problem Efficiently learning and analyzing data streams with limited memory.
method Introduces Weight-Median Sketch for compressed linear classifier learning over data streams.
result Memory-limited execution of various analyses over streams, including feature selection and mutual information estimation.
Dynamic Model Tree improves online learning for evolving data streams.
problem Effective and transparent machine learning on data streams is challenging.
method Revisit Model Trees for data stream applications, introducing Dynamic Model Tree.
result Dynamic Model Tree reduces the number of splits and outperforms state-of-the-art models.
New framework improves fraud prediction with incremental data balancing for massive data streams.
problem Class imbalance problem in massive imbalanced data streams.
method Incremental data balancing framework using Racing Algorithm for automated balancing and Random Forest for classification.
result Better results than Batch mode on European Credit Card dataset.
Survey examines big data's role in network design.
problem Designing robust networks with refined performance and intelligent features.
method Integrates big data analytics with network control/traffic layers.
result Big data analytics can improve network design.
Adapts DPMM for fast streaming data clustering.
problem Clustering streaming data with time-dependent statistics.
method Adapts DPMM and sampling-based inference for online clustering.
result Obtains state-of-the-art results in speed and accuracy.
CROC identifies the earliest-changing stream as the root cause in multi-stream data.
problem Distribution-free root cause analysis in multi-stream data with unknown distributional changes.
method Conformal p-values and finite-sample valid confidence sets.
result CROC efficiently isolates the root cause under minimal assumptions.
DSSCN improves lifelong learning of non-stationary data streams through adaptive network construction.
problem Lifelong learning of non-stationary data streams with efficient and adaptive models.
method Deep stacked stochastic configuration network (DSSCN) with self-constructing deep stacked network structure and adaptive hidden unit parameters.
result DSSCN outperforms existing data stream algorithms in continual learning of non-stationary data streams.