CURIE uses cellular automata to detect concept drift in data streams.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We adapted the Covertype data set for unsupervised learning.
CICLAD efficiently mines frequent closed itemsets from data streams with minimal memory usage.
The aim of process discovery, originating from the area of process mining, is to discover a process model based on business process execution data. A majority of process discovery techniques relies on an event log as an input. An event log is a static source of historical data capturing the execution of a business proc…
OLBoost improves online decision tree performance without increasing memory or time costs.
New algorithm detects and handles concept drift in data streams.
In data stream mining, predictive models typically suffer drops in predictive performance due to concept drift. As enough data representing the new concept must be collected for the new concept to be well learnt, the predictive performance of existing models usually takes some time to recover from concept drift. To spe…
Scikit-multiflow is a multi-output/multi-label and stream data mining framework for the Python programming language. Conceived to serve as a platform to encourage democratization of stream learning research, it provides multiple state of the art methods for stream learning, stream generators and evaluators. scikit-mult…
StreaMRAK improves KRR for streaming data.
Paper proposes a framework to detect adversarial concept drifts under poisoning attacks.
A framework for evaluating and benchmarking concept drift detection methods
New method improves prediction accuracy in business process mining by handling concept drift.
A number of recent emerging applications call for studying data streams, potentially infinite flows of information updated in real-time. When multiple co-evolving data streams are observed, an important task is to determine how these streams depend on each other, accounting for dynamic dependence patterns without impos…
New method combines deep learning and streaming learning for better incremental learning.
Stream mining poses unique challenges to machine learning: predictive models are required to be scalable, incrementally trainable, must remain bounded in size (even when the data stream is arbitrarily long), and be nonparametric in order to achieve high accuracy even in complex and dynamic environments. Moreover, the l…
Learning from data streams is an increasingly important topic in data mining, machine learning, and artificial intelligence in general. A major focus in the data stream literature is on designing methods that can deal with concept drift, a challenge where the generating distribution changes over time. A general assumpt…
Dynamic submodular maximization with consistency constraints.
Novel algorithm SAODE improves high-dimensional stream classification in seasonal data.
Tensor decompositions are used in various data mining applications from social network to medical applications and are extremely useful in discovering latent structures or concepts in the data. Many real-world applications are dynamic in nature and so are their data. To deal with this dynamic nature of data, there exis…
Paper addresses challenges in benchmarking stream learning algorithms with real-world data.
Automated process discovery is a class of process mining methods that allow analysts to extract business process models from event logs. Traditional process discovery methods extract process models from a snapshot of an event log stored in its entirety. In some scenarios, however, events keep coming with a high arrival…
Proposes a new online learning strategy for multi-target regression in data streams.
We propose scalable methods to execute counting queries in machine learning applications. To achieve memory and computational efficiency, we abstract counting queries and their context such that the counts can be aggregated as a stream. We demonstrate performance and scalability of the resulting approach on random quer…
As data streams become more prevalent, the necessity for online algorithms that mine this transient and dynamic data becomes clearer. Multi-label data stream classification is a supervised learning problem where each instance in the data stream is classified into one or more pre-defined sets of labels. Many methods hav…
Method detects interactions for better CTR prediction.
Operating in a dynamic real world environment requires a forward thinking and adversarial aware design for classifiers, beyond fitting the model to the training data. In such scenarios, it is necessary to make classifiers - a) harder to evade, b) easier to detect changes in the data distribution over time, and c) be ab…
FASE-AL uses active learning to reduce labeling costs for data streams.
Correlated anomaly detection (CAD) from streaming data is a type of group anomaly detection and an essential task in useful real-time data mining applications like botnet detection, financial event detection, industrial process monitor, etc. The primary approach for this type of detection in previous researches is base…
Many tasks in machine learning and data mining, such as data diversification, non-parametric learning, kernel machines, clustering etc., require extracting a small but representative summary from a massive dataset. Often, such problems can be posed as maximizing a submodular set function subject to a cardinality constr…
Decision tree classifiers are a widely used tool in data stream mining. The use of confidence intervals to estimate the gain associated with each split leads to very effective methods, like the popular Hoeffding tree algorithm. From a statistical viewpoint, the analysis of decision tree classifiers in a streaming setti…
TDX predicts evolving distributions from streaming data.
SAFAVI framework detects anomalies without a single best solution.
ESRF reduces ARF ensemble size without sacrificing accuracy.
Distributed, online data mining systems have emerged as a result of applications requiring analysis of large amounts of correlated and high-dimensional data produced by multiple distributed data sources. We propose a distributed online data classification framework where data is gathered by distributed data sources and…
This study designs a financial risk control platform using big data and machine learning.
Game-theoretic analysis of mining gaps in blockchain systems.
Alpha-GPT mines new trading signals with human-AI interaction.
The paper explores how mining costs, rewards, and blockchain security are interconnected.
Data preprocessing improves data quality for robust data mining.
This paper categorizes and analyzes existing outlying aspect mining methods.
Game theory shows miners' hardware improvements don't centralize mining.
DataLearner simplifies data mining on Android devices.
A new method classifies multiple correlated data streams simultaneously.
A new method for estimating SW from streaming data.
This paper analyzes the profitability of selfish mining on blockchain, considering the risk of ruin.
Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft wares are available, but data mining in agricultural soil datasets is a relatively…
This paper tracks coin circulation in Bitcoin to identify miners and analyze mining pool structures.
We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, an…