A new method for finding significant patterns in stratified data.
problem Statistical challenges in finding enriched itemsets in stratified data.
method Proposes a strategy and efficient algorithm for significant pattern mining in the presence of categorical covariates.
result Efficiently corrects for multiple testing in pattern mining while retaining statistical power.
FSR efficiently discovers significant patterns with few resampled datasets.
problem Mining significant patterns in transactional data, especially subgroups.
method FSR uses resampling to bound the supremum deviation of quality statistics, providing rigorous guarantees on false discoveries.
result FSR effectively discovers significant subgroups with a small number of resampled datasets.
New approach finds statistically significant patterns from databases.
problem Difficulty in statistically significant pattern discovery due to selection bias.
method Selective inference framework applied to pattern mining problems.
result Effective in finding statistically significant patterns from databases.
Discovers discriminative patterns in two-class datasets.
problem Discovering patterns that occur more frequently in one class than the other.
method Proposes SSDPS algorithm with an original enumeration strategy exploiting anti-monotonicity.
result SSDPS outperforms other algorithms in terms of efficiency and pattern generation.
We present a novel algorithm, Westfall-Young light, for detecting patterns, such as itemsets and subgraphs, which are statistically significantly enriched in one of two classes. Our method corrects rigorously for multiple hypothesis testing and correlations between patterns through the Westfall-Young permutation proced…
Data mining enhances a heuristic for the Minimum Latency Problem.
problem Finding optimal solutions for the Minimum Latency Problem efficiently.
method Combining GRASP with data mining to find frequent patterns in high-quality solutions.
result Improved solution quality and reduced computational time compared to existing methods.
MCRapper efficiently computes patterns in data using Monte-Carlo Rademacher Averages.
problem Finding statistically significant patterns in data with limited samples.
method Monte-Carlo Empirical Rademacher Averages (MCERA) for poset families.
result MCRapper provides upper bounds to the discrepancy of functions, enabling efficient pattern mining.
Safe Pattern Pruning reduces pattern explosion in predictive pattern mining.
problem Exponential growth of patterns in structured data.
method Safe Pattern Pruning (SPP) method.
result Effective model building in practical data analysis.
Randomization helps verify if data mining results are due to inherent patterns.
problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.
Safe pattern pruning finds predictive patterns efficiently.
problem Finding optimal predictive patterns in databases.
method Safe Pattern Pruning (SPP) method for predictive pattern mining.
result SPP method efficiently finds superset of all needed predictive patterns.
LetSIP learns relevant patterns for user interests in data mining.
problem Redundancy in pattern mining makes it hard for analysts to identify relevant patterns.
method Combines pattern sampling with interactive data mining, using user feedback to learn sampling distribution.
result Favourable trade-offs in quality-diversity and exploitation-exploration compared to existing methods.
A new model for efficient sequential pattern mining without specific encoding schemes.
problem Mining relevant sequential patterns efficiently and effectively.
method Subsequence interleaving model based on probabilistic sequence database.
result Efficient inference through submodular optimization, resulting in low spuriousness and redundancy.
A new framework for mining high utility patterns in interval-based sequences.
problem Mining patterns in events that persist over varying time intervals and considering event utility.
method Integrates utility into interval-based sequences and proposes HUIPMiner algorithm with pruning strategy.
result HUIPMiner efficiently finds high utility patterns in real datasets.
Proposes SCR-Apriori for efficient mining of SCR-patterns.
problem Mining high-quality `Set of Contrasting Rules'-pattern (SCR-pattern) efficiently.
method Integrates SCR-pattern structure into Apriori algorithm to prune search space.
result Significantly reduces computational cost compared to state-of-the-art.
The study compares ASP encodings for sequential pattern mining tasks.
problem Efficiency of Answer Set Programming (ASP) encodings for sequential pattern mining.
method Two representations of embeddings (fill-gaps vs skip-gaps) and various types of patterns were tested.
result Fill-gaps strategy is more efficient on real problems due to lower memory consumption.
FedSLIM optimizes compact pattern models across distributed databases without sharing raw data.
problem Privacy-preserving federated descriptive analytics for data silos.
method Federated MDL-based framework using SLIM principle.
result FedSLIM variants preserve high-quality compression structure and recover globally informative patterns.
Paper mines rank data patterns from rankings.
problem Mining rank data patterns from rankings.
method Proposes algorithms for frequent rankings and dependencies.
result Experimental validation of algorithms on synthetic and real data.
New method for faster TPM from multivariate time series.
problem Mining predictive complex temporal patterns from multivariate time series.
method Fast Temporal Pattern Mining with Extended Vertical Lists.
result Significantly faster performance than previous algorithm.
SDSM extracts statistically significant sub-trajectories from large trajectory datasets.
problem Discerning moving patterns that are more characteristic of one group of trajectories than another.
method Statistically Discriminative Sub-trajectory Mining (SDSM) method using tree representation and permutation-based statistical inference.
result SDSM efficiently extracts statistically significant sub-trajectories from massive trajectory datasets.
R-GPM enables efficient graph pattern mining through user-defined relations.
problem Efficient graph pattern mining through user-defined relations.
method Parallel computing framework with MCMC sampling algorithm and optimizations.
result Efficient estimators for graph pattern statistics with up to 3-orders-of-magnitude computational cost reduction.
The problem of multiple hypothesis testing arises when there are more than one hypothesis to be tested simultaneously for statistical significance. This is a very common situation in many data mining applications. For instance, assessing simultaneously the significance of all frequent itemsets of a single dataset entai…
Flexics samples patterns with guarantees, addressing flexibility and accuracy issues.
problem Pattern explosion and limited sampling accuracy with existing methods.
method Leverages SAT sampling and pattern mining algorithms to support flexible quality measures and constraints.
result Flexics provides strong guarantees on sampling accuracy while being flexible and efficient.
This study proposes a deep learning framework using ResNeXt for efficient financial data mining.
problem Complex financial data with high dimensionality, nonlinearity, and task correlations.
method Introduces ResNeXt into multi-task learning framework for efficient feature extraction and task collaboration.
result Significantly improved performance in classification and regression tasks on S&P 500 data.
Deep learning applied to biological data mining.
problem Mining complex biological data from diverse sources.
method Artificial neural networks, deep learning architectures.
result Deep learning techniques improve pattern recognition in biological data.
Study discovers semantic patterns in consumer sentiment on social media.
problem Identifying semantic patterns in consumer sentiment expressed in tweets.
method Cosine similarity, K-means clustering, and Latent Dirichlet Allocation (LDA) were used.
result Identified latent topics representing consumer opinions on social media.
SGE learns symbolic node representations from relational data.
problem Mining insights from complex, real-world systems.
method SGE uses frequent pattern mining on a node's neighborhood to learn symbolic node representations.
result SGE outperforms shallow node embedding methods on a venue classification task.
Skopus discovers top-k sequential patterns with leverage.
problem Discovering top-k sequential patterns with leverage.
method Combines novel expected support definition with SkOPUS algorithm.
result Exact discovery of top-k sequential patterns with leverage.
Clusters energy usage patterns from smart meters.
problem Identify and group similar energy usage profiles.
method Clustering time-series data from smart meters.
result Accurate grouping of similar energy usage patterns.
This study identifies RwD crash patterns on rural two-lane highways under different lighting conditions.
problem Insufficient investigation of RwD crashes under varying lighting conditions.
method Data mining using association rules mining (ARM) on crash database.
result Interesting crash patterns and risk factors identified under different lighting conditions.
Method extracts taint flows to classify Bitcoin mining pools.
problem Understanding pseudonymous Bitcoin actors and their transactions.
method Taint analysis and graph embedding methods applied to taint flows.
result Taint flows from the same period show high similarity.
Paper detects common subtrees with identical labels in trees.
problem Finding common subtrees with identical label distribution in tree data.
method Developed an algorithm for tree isomorphism and a new compression scheme for trees.
result The method efficiently finds and compresses common subtrees with identical labels.
Paper proposes a method to enhance graph classification for neurological disorders using multiple side views.
problem Discerning subgraph patterns from graph data alone is insufficient for disease diagnosis.
method Developed a novel approach to select optimal subgraph features by integrating multiple side views.
result Subgraph patterns selected using multiple side views improve graph classification for neurological disorders.
The problem of finding itemsets that are statistically significantly enriched in a class of transactions is complicated by the need to correct for multiple hypothesis testing. Pruning untestable hypotheses was recently proposed as a strategy for this task of significant itemset mining. It was shown to lead to greater s…
FLEXI optimizes binning for better subgroup discovery in numerical and ordinal attributes.
problem Mining high quality subgroups from numerical attributes is challenging.
method FLEXI uses optimal binning to find high quality binary features for both numeric and ordinal attributes.
result FLEXI outperforms state of the art with up to 25 times improvement in subgroup quality.
NegPSpan extracts negative sequential patterns efficiently.
problem Mining negative sequential patterns in sequence datasets.
method PrefixSpan depth-first scheme with maxgap constraints.
result NegPSpan extracts meaningful negative sequential patterns.
The paper advances OA-biclustering for multi-mode community detection in social networks.
problem Mining meaningful patterns in multi-mode networks for community detection.
method Object-attribute biclustering (OA-biclustering) for 2-mode networks, extended to 3- and 4-mode networks.
result OA-biclusters are suitable for community detection in multi-mode cases, even with unknown number of corresponding n-cliques. DM4OG workshop tackles data mining in oil and gas industry.
problem Data growth challenges in oil and gas industry.
method Data mining techniques and machine learning.
result Effective data insight for decision making.
The paper improves itemset quality assessment by incorporating background knowledge.
problem Assessing the quality of discovered itemsets is challenging due to many patterns being explainable by background knowledge.
method The authors introduce a maximum entropy approach to efficiently infuse additional background knowledge such as row margins, lazarus counts, and bounds of ones.
result More sophisticated models that incorporate background knowledge fit the data better and improve frequency prediction of itemsets.
Paper presents RSVD for better recommender system performance.
problem Improving recommender system performance.
method Regularized SVD (RSVD) with efficient algorithm and theoretical analysis.
result RSVD outperforms SVD in recommender systems.
The paper proposes a novel approach to music analysis using text mining techniques.
problem Analyzing musical documents using traditional text mining methods.
method Developed a Naive Dictionary of 'muselets' (musical words) of uniform length.
result Demonstrated reasonable topic modeling and pattern recognition results with a simplified dictionary.
Game theory shows miners' hardware improvements don't centralize mining.
problem Decentralization of cryptocurrency mining.
method Game-theoretical model of mining efficiency and competition.
result Advancements in mining hardware efficiency do not lead to centralization.
A new method improves maximum margin criterion for better pattern analysis.
problem Handling high dimensionality and large data in pattern analysis.
method Introducing an improved maximum margin criterion (MMC) and its variants.
result Experimental results show the MMC methods are effective in complex scenarios.
Paper uses data to teach smart homes to save energy.
problem Achieving energy savings in smart homes.
method Frequent sequential pattern mining algorithm for real-life smart home data.
result The recommender system reduces energy consumption effectively.
New biclustering algorithms for microarray data using Formal Concept Analysis.
problem Uncovering patterns in gene expression data matrices.
method Formal Concept Analysis and Association Rules.
result Promising results from proposed biclustering algorithms.
Random Intersection Chains selects important interactions from categorical features.
problem Heavy computational burden in considering all interactions for categorical features.
method Randomly generates chains of intersections, estimates and selects frequent patterns.
result Selected patterns are the most frequent in the data set.
FIBS extracts relevant features from IBTSs for classification.
problem Classifying interval-based temporal sequences (IBTSs) using common algorithms is challenging.
method FIBS extracts features from IBTSs based on relative frequency and temporal relations, incorporating a filter-based selection strategy to avoid irrelevant features.
result FIBS effectively represents IBTSs for classification algorithms, providing similar or better accuracy compared to state-of-the-art competitors.
QuantaAlpha uses evolutionary algorithms to mine financial alpha robustly across market distributions.
problem Challenges in alpha mining due to market noise and regime shifts.
method Evolutionary framework treating each mining run as a trajectory, mutation, crossover, targeted revision, and reuse of effective patterns.
result Consistent gains over strong baselines and prior systems, achieving high IC and ARR.
Study shows Bitcoin security tied to mining rewards and prices.
problem Understanding Bitcoin security's dependency on market outcomes.
method Used ARDL approach with daily blockchain and Bitcoin data from 2014-2019.
result Bitcoin security outcomes linked to Bitcoin price and mining rewards.