Systematic review of ML explainability in process mining.
problem Understanding the black-box nature of ML models in process mining.
method Systematic literature review using PRISMA framework.
result Identification of key trends and challenges in interpretability.
Novel process model for metabolomics data analysis.
problem Analyzing complex metabolomics data.
method Data-driven and hypothesis-driven data mining approaches using various techniques.
result Demonstrated applicability and strengths of MeKDDaM model.
In January 3, 2009, Satoshi Nakamoto gave rise to the "Bitcoin Block Chain" creating the first block of the chain hashing on his computers central processing unit (CPU). Since then, the hash calculations to mine Bitcoin have been getting more and more complex, and consequently the mining hardware evolved to adapt to th…
This paper evaluates conformance measures in process mining using conformance propositions.
problem Lack of formal definition and evaluation of conformance measures in process mining.
method Formulated 21 conformance propositions to evaluate existing measures.
result Identified challenges and requirements for conformance measures in process mining.
Process mining is a research field focused on the analysis of event data with the aim of extracting insights in processes. Applying process mining techniques on data from smart home environments has the potential to provide valuable insights in (un)healthy habits and to contribute to ambient assisted living solutions. …
The paper uses machine learning to find causal rules from business process logs.
problem Discovering causal relationships in business process logs.
method Action rule mining followed by causal machine learning (uplift trees).
result Identifies treatments with high causal effect on outcomes.
Study proposes a novel local explanation method for deep learning classifiers in process mining.
problem Lack of interpretability in deep learning models for process mining.
method Defines local regions using latent space representations and visualizes explanations.
result Deep learning classifier achieves high performance and local explanations increase user trust.
New method improves prediction accuracy in business process mining by handling concept drift.
problem Improving prediction quality in business process mining affected by concept drift.
method Systematically analyzed and compared different data selection strategies for retraining machine learning models.
result Improved accuracy from 0.5400 to 0.7010 with concept drift handling.
Graph neural networks detect anomalies in object-centric business processes.
problem Detecting anomalies in graph-like business processes.
method Graph convolutional autoencoder architecture for anomaly detection.
result Promising performance in detecting anomalies at the activity type and attributes level.
Software automates metabolomics data analysis for reproducible results.
problem Automating reproducible metabolomics data analysis.
method Object-oriented software engineering, Java, XML database, GUI, version control system.
result MeKDDaM-SAGA successfully guides metabolomics applications.
Analyzes quantitative finance papers from arXiv using text mining and NLP.
problem Understanding trends and insights in quantitative finance research.
method Text mining, natural language processing, topic modeling.
result Identified most cited researchers and journals in quantitative finance.
DataLearner simplifies data mining on Android devices.
problem Lack of general-purpose data-mining tools for mobile devices.
method Augments Weka engine with Charles Sturt University algorithms, providing 40 mining algorithms.
result Delivers classification accuracy similar to PCs/laptops with acceptable speed and battery life.
GRM uses graph neural networks to score process activity relevance.
problem Improving business processes with performance measures.
method Graph Relevance Miner (GRM) based on graph neural networks.
result Quantitatively evaluated relevance scores with four datasets.
Peer-reviewed research and mined data predict stock returns similarly.
problem Predicting stock returns using research quality.
method Cross-sectional analysis of 29,000 accounting ratios with t-statistics > 2.0.
result Post-sample performance is largely independent of whether the predictor is peer-reviewed or mined.
The paper predicts workload using process mining and neural networks.
problem Predicting workload in business processes.
method Process mining logs, combined with recurrent neural networks.
result Workload prediction achieved with an MAPE score of 19%.
Recurrent neural networks improve process instance classification.
problem Classifying ongoing process instances based on activities.
method Applied recurrent neural networks, specifically GRU, to classify business process instances.
result GRU outperforms LSTM in training time with similar accuracy.
Improves event prediction in complex processes using Petri nets and deep learning.
problem Predicting the next event in complex processes given a state.
method Enhanced Petri net model with time decay functions and deep learning.
result Significant performance improvements over state-of-the-art methods.
Safe Pattern Pruning reduces pattern explosion in predictive pattern mining.
problem Exponential growth of patterns in structured data.
method Safe Pattern Pruning (SPP) method.
result Effective model building in practical data analysis.
TLRS improves predictive power of mined formulaic alpha factors.
problem Sparse rewards in RL for mining formulaic alpha factors.
method Trajectory-level Reward Shaping (TLRS) with reward centering.
result TLRS boosts predictive power by 9.29% over existing methods.
The paper uses GIS data to predict urban sprawl.
problem Overgrowth and expansion of low-density areas with car dependency and segregation.
method Data mining algorithms (Apriori, J4.8) adapted for geospatial analysis using ArcGIS.
result Prototype spatial decision support system (SDSS) predicts urban sprawl and estimates impact variables.
CONDA-PM framework helps analyze concept drift in business processes.
problem Analyzing changes in business processes over time.
method Systematic Literature Review and framework development.
result Highlights areas needing research to complement existing efforts.
The purpose of this study is to estimate the production function and examine the structure of production in the mining sector of Iran. Several studies have already been conducted in estimating production functions of various economic sectors; however, less attention has been paid to mining sectors. After examining the …
Paper proposes active learning for mining conflict dynamics from textual data.
problem Lack of detailed information on conflict dynamics in existing datasets.
method Active learning with a large, encoder-only language model.
result Performance similar to human coding with reduced human annotation.
A new tool, matrix profile, finds all pair similarities in time series data.
problem Finding all pair similarities in time series data.
method Near universal time series data mining tool called matrix profile.
result Matrix profile solves the all-pairs-similarity-search problem for time series subsequences.
The aim of process discovery, originating from the area of process mining, is to discover a process model based on business process execution data. A majority of process discovery techniques relies on an event log as an input. An event log is a static source of historical data capturing the execution of a business proc…
RiskMiner discovers formulaic alphas using MCTS for better performance.
problem Mining formulaic alphas without considering structural information and alpha correlations.
method Formulates alpha mining as an MDP and solves it with a risk-seeking MCTS.
result Our method outperforms state-of-the-art benchmarks and achieves the most profitable results.
Framework extracts symptoms from EHRs for rapid disease outbreak detection.
problem Extracting relevant data from unstructured medical texts.
method Conformal active learning for efficient data mining.
result Framework achieves strong performance with minimal manual labeling.
Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that patterns are selected from extremely large number of candidates in databases. In …
Proposes MLDP for modeling multilinear data.
problem Handling data with interactions from multiple factors.
method Combines Dirichlet processes with multilinear factor analysis.
result Achieved state-of-the-art performance on real-world data.
This paper explores NLP techniques for insurance, detailing methods and applications.
problem Extracting value from insurance reports using complex text data.
method Detailed explanation of NLP methods and their implementation in insurance.
result Enhanced risk monitoring and policyholder benefits through NLP.
Tool uses text mining to define innovative tech fields from abstracts.
problem Defining scope in dynamic, multidisciplinary tech projects.
method Text mining of Elsevier's Scopus abstracts.
result Tool provides crucial information for tech field definition.
FastForest boosts Random Forest speed by 24%.
problem Efficiency in processing speed for Random Forest.
method Subsample Aggregating, Logarithmic Split-Point Sampling, Dynamic Restricted Subspacing.
result Average 24% increase in processing speed with accuracy maintained.
New method expands seed genes to functionally related clusters.
problem Discovering functionally related genes lacking GO terms.
method Semi-supervised learning with positive and unlabeled examples.
result LPU approaches significantly outperform existing methods.
Study proposes explainable analytics for manufacturing process planning.
problem Improving data-driven decision-making in manufacturing.
method Combines process mining, machine learning, and XAI. Uses deep learning for prediction and Shapley values/ICE plots for explanations.
result Enhanced decision-making capabilities through local post-hoc explanations.
Develops a non-parametric Dirichlet process method for probabilistic biclustering.
problem Challenges in finding biclusters with strong co-occurrence in rows and columns.
method Dual Dirichlet process mixture models for row and column clustering, with cluster number determined by data.
result Improves bicluster extraction in text mining and gene expression analysis.
We give an explicit definition of decentralization and show you that decentralization is almost impossible for the current stage and Bitcoin is the first truly noncentralized currency in the currency history. We propose a new framework of noncentralized cryptocurrency system with an assumption of the existence of a wea…
Automated process discovery is a class of process mining methods that allow analysts to extract business process models from event logs. Traditional process discovery methods extract process models from a snapshot of an event log stored in its entirety. In some scenarios, however, events keep coming with a high arrival…
Interdisciplinary comparison of sequence modeling methods for next-element prediction.
problem Comparing sequence modeling methods across different fields.
method Experimental evaluation of four real-life sequence datasets using machine learning, process mining, and grammar inference techniques.
result Machine learning techniques outperform interpretability-focused methods in next-element prediction accuracy.
This study uses NLP to detect financial risks from documents.
problem Detecting and predicting financial risks in documents.
method NLP model design, text preprocessing, feature extraction, machine learning.
result NLP model effectively identifies and predicts financial risks.
Optimising black-box functions is important in many disciplines, such as tuning machine learning models, robotics, finance and mining exploration. Bayesian optimisation is a state-of-the-art technique for the global optimisation of black-box functions which are expensive to evaluate. At the core of this approach is a G…
OLBoost improves online decision tree performance without increasing memory or time costs.
problem Improving predictive performance in online decision trees without high memory or time costs.
method OLBoost applies boosting to small regions of the instances space within online decision tree algorithms.
result OLBoost can significantly improve online learning decision tree performance without increasing tree size.
We adapted the Covertype data set for unsupervised learning.
problem Lack of suitable unsupervised learning data sets.
method Transformed the Covertype data set into the Wilderness Area data set.
result The Wilderness Area data set is more suitable for unsupervised learning.
Game-theoretic analysis of mining gaps in blockchain systems.
problem Strategic mining behavior and its impact on blockchain stability.
method Game-theoretic model and Nash equilibrium analysis.
result Mining gaps can destabilize blockchain systems, especially with decreasing block rewards.
Randomization helps verify if data mining results are due to inherent patterns.
problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
Paper proposes a new framework to mine synergistic formulaic alphas for better stock trend forecasting.
problem Mining alphas separately ignores their combined performance, leading to suboptimal models.
method Proposes a reinforcement learning-based framework that optimizes the mining of synergistic formulaic alpha sets.
result Demonstrates higher returns in stock trend forecasting compared to previous approaches.
The process of exploring and exploiting Oil and Gas (O&G) generates a lot of data that can bring more efficiency to the industry. The opportunities for using data mining techniques in the "digital oil-field" remain largely unexplored or uncharted. With the high rate of data expansion, companies are scrambling to develo…
Alpha-GPT mines new trading signals with human-AI interaction.
problem Mining new alphas for effective trading signals.
method Human-AI interaction and prompt engineering algorithmic framework.
result Demonstrates Alpha-GPT's effectiveness in generating creative, insightful, and effective alphas.