The paper tackles sample complexity in high-dimensional data, focusing on correlation mining.
problem Understanding reliable inference in variable-rich, sample-starved data.
method Develops a unified statistical framework to quantify sample complexity for various inferential tasks.
result Illustrates high-dimensional learning rates and sample complexity for correlation mining.
Data mining techniques on the biological analysis are spreading for most of the areas including the health care and medical information. We have applied the data mining techniques, such as KNN, SVM, MLP or decision trees over a unique dataset, which is collected from 16,380 analysis results for a year. Furthermore we h…
The paper explores how mining costs, rewards, and blockchain security are interconnected.
problem Understanding the interdependencies between mining costs, mining rewards, and blockchain security.
method Theoretical derivation and empirical analysis using daily crypto market data and autoregressive distributed lag approach.
result Cryptocurrency price and mining rewards are intrinsically linked to blockchain security outcomes.
AlphaEval evaluates alpha mining models efficiently and comprehensively.
problem Lack of systematic evaluation for alpha mining models.
method Unified, parallelizable evaluation framework assessing predictive power, stability, robustness, financial logic, and diversity.
result AlphaEval achieves evaluation consistency comparable to comprehensive backtesting, providing more comprehensive insights and higher efficiency.
New biclustering algorithms for microarray data using Formal Concept Analysis.
problem Uncovering patterns in gene expression data matrices.
method Formal Concept Analysis and Association Rules.
result Promising results from proposed biclustering algorithms.
Randomization helps verify if data mining results are due to inherent patterns.
problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.
UDS measures multivariate correlations across any dimensionality.
problem Discovering significant correlations in multi-dimensional data.
method UDS based on cumulative entropy, normalized for comparison across subspaces.
result UDS efficiently captures both linear and non-linear correlations.
This paper offers a new algebraic perspective of GCCA using subspace intersection.
problem Finding common variables across multiple feature representations.
method Subspace intersection approach based on a (bi-)linear generative model.
result GCCA is equivalent to subspace intersection, with conditions for identifiable common subspace.
We present a novel algorithm, Westfall-Young light, for detecting patterns, such as itemsets and subgraphs, which are statistically significantly enriched in one of two classes. Our method corrects rigorously for multiple hypothesis testing and correlations between patterns through the Westfall-Young permutation proced…
This paper improves image super-resolution by integrating cross-scale non-local attention.
problem Improving image super-resolution by leveraging long-range and cross-scale feature correlations.
method Proposes a Cross-Scale Non-Local (CS-NL) attention module integrated into a recurrent neural network.
result Significantly improved performance on SISR benchmarks.
AlphaSAGE mines diverse alphas via GFlowNets, overcoming RL issues.
problem Reward sparsity, inadequate sequential representations, and single optimal mode issues in RL for alphas.
method Structure-aware encoder (RGCN), GFlowNets, dense reward structure.
result Empirically outperforms existing baselines in mining diverse alphas.
The paper uses machine learning to find causal rules from business process logs.
problem Discovering causal relationships in business process logs.
method Action rule mining followed by causal machine learning (uplift trees).
result Identifies treatments with high causal effect on outcomes.
A new framework forecasts stock trends by mining shared information from concepts.
problem Forecasting stock trends using static concept information limits accuracy.
method Proposes a graph-based framework that mines concept-oriented shared information from both predefined and hidden concepts.
result Improves stock trend forecasting performance through dynamic concept relevance and hidden concept information.
Paper proposes a new REINFORCE algorithm for mining formulaic alpha factors with reduced variance.
problem Mining formulaic alpha factors with interpretability and robustness in volatile markets.
method Developed a novel REINFORCE algorithm with a dedicated baseline and reward shaping.
result Boosts correlation with returns by 3.83% and enhances excess returns compared to existing methods.
RiskMiner discovers formulaic alphas using MCTS for better performance.
problem Mining formulaic alphas without considering structural information and alpha correlations.
method Formulates alpha mining as an MDP and solves it with a risk-seeking MCTS.
result Our method outperforms state-of-the-art benchmarks and achieves the most profitable results.
Improves data efficiency for estimating mutual information.
problem Estimating mutual information from limited data.
method Developed a Data-Efficient MINE Estimator (DEMINE) and Meta-DEMINE.
result Significantly improved data efficiency in estimating mutual information.
This study proposes a deep learning framework using ResNeXt for efficient financial data mining.
problem Complex financial data with high dimensionality, nonlinearity, and task correlations.
method Introduces ResNeXt into multi-task learning framework for efficient feature extraction and task collaboration.
result Significantly improved performance in classification and regression tasks on S&P 500 data.
ProxiModel extracts high-quality news events from news corpora.
problem Mining high-quality structured event knowledge from noisy news data.
method ProxiModel uses a proximity-network to model event correlation within and across news corpora.
result ProxiModel efficiently and effectively extracts high-quality event descriptors and attributes.
New metrics unify and generalize popular distances for data mining.
problem Unified and generalized distances for data mining.
method Introducing new metrics on sets, vectors, and functions.
result New metrics outperform traditional ones in real-valued and structured data.
IFGAN uses feature-specific GANs for missing value imputation.
problem Missing value imputation in data mining.
method Feature-specific Generative Adversarial Networks (GAN).
result IFGAN outperforms state-of-the-art algorithms in various missing conditions.
Feature selection has attracted significant attention in data mining and machine learning in the past decades. Many existing feature selection methods eliminate redundancy by measuring pairwise inter-correlation of features, whereas the complementariness of features and higher inter-correlation among more than two feat…
Proposes MLDP for modeling multilinear data.
problem Handling data with interactions from multiple factors.
method Combines Dirichlet processes with multilinear factor analysis.
result Achieved state-of-the-art performance on real-world data.
Discover novel multivariate relationships in time series data.
problem Capturing novel relationships between time series in complex systems.
method Introducing multipoles as linear relationships among more than two time series, identifying them as cliques of negative correlations in a correlation network.
result Almost all multipoles can be efficiently found using a clique-enumeration approach.
FactorMiner discovers financial alpha factors with low redundancy.
problem Finding novel financial alpha factors in a vast search space.
method Modular Skill Architecture and Experience Memory to distill and guide exploration.
result FactorMiner constructs a diverse library of high-quality factors with competitive performance.
A new imputation model for clinical data captures both cross-sectional and temporal correlations.
problem Missing values in multivariable time series clinical data.
method Integrates Gaussian processes with mixture models and individualized mixing weights.
result The proposed model provides more accurate imputation than benchmarks on real-world and synthetic datasets.
This paper deals with the binary classification task when the target class has the lower probability of occurrence. In such situation, it is not possible to build a powerful classifier by using standard methods such as logistic regression, classification tree, discriminant analysis, etc. To overcome this short-coming o…
Game-theoretic analysis of mining gaps in blockchain systems.
problem Strategic mining behavior and its impact on blockchain stability.
method Game-theoretic model and Nash equilibrium analysis.
result Mining gaps can destabilize blockchain systems, especially with decreasing block rewards.
A neural network method to measure correlation among multiple variables.
problem Measuring correlation among multiple variables.
method Designing a neural network function, calculating differences, optimizing parameters, and using POET algorithm.
result The method can improve neural network performance and address issues like overfitting.
Predicts coal mine seismic events up to 8 hours in advance.
problem Early detection of dangerous seismic events in coal mines.
method Ensemble machine learning models trained on various features.
result Best model achieved 0.939 AUC.
New methods improve anomaly detection from large streaming data.
problem Anomaly detection in large streaming data fails with current methods.
method Two novel randomized algorithms (rPS and gPS) for better detection of correlated anomalies.
result High and balanced recall and estimated accuracy for anomaly detection.
Alpha-GPT mines new trading signals with human-AI interaction.
problem Mining new alphas for effective trading signals.
method Human-AI interaction and prompt engineering algorithmic framework.
result Demonstrates Alpha-GPT's effectiveness in generating creative, insightful, and effective alphas.
This study uses deep learning to analyze stock market sentiment from financial forums.
problem Improving stock market prediction accuracy through emotional analysis.
method Crawling financial forum data, training Bert model on financial corpus, using MIC for comparison.
result BERT model's emotional analysis of financial texts correlates with stock market fluctuations.
Data preprocessing improves data quality for robust data mining.
problem Noisy and incomplete data hinders data mining models.
method Overview of data cleaning, transformation, and preprocessing methods.
result Preprocessing significantly affects data mining model performance.
This paper categorizes and analyzes existing outlying aspect mining methods.
problem Finding unique features in data objects that differ from others.
method Grouping and analyzing existing outlying aspect mining approaches in three categories.
result Comparison of strengths, weaknesses, and time complexities of different techniques.
Game theory shows miners' hardware improvements don't centralize mining.
problem Decentralization of cryptocurrency mining.
method Game-theoretical model of mining efficiency and competition.
result Advancements in mining hardware efficiency do not lead to centralization.
DataLearner simplifies data mining on Android devices.
problem Lack of general-purpose data-mining tools for mobile devices.
method Augments Weka engine with Charles Sturt University algorithms, providing 40 mining algorithms.
result Delivers classification accuracy similar to PCs/laptops with acceptable speed and battery life.
Agent-based model simulates Bitcoin mining economics.
problem Modeling the economics of Bitcoin mining.
method Agent-based artificial market model of Bitcoin mining and transactions.
result Model reproduces key financial features of Bitcoin.
This paper analyzes the profitability of selfish mining on blockchain, considering the risk of ruin.
problem The profitability of selfish mining on blockchain, considering the risk of ruin.
method Formulated a stochastic model and used tools from applied probability and analysis to determine expected profit.
result Explicit expressions for expected profit under different scenarios were derived, identifying conditions for selfish mining as a strategic advantage.
MINE uses neural networks to estimate mutual information efficiently.
problem Estimating mutual information between high-dimensional variables.
method Gradient descent over neural networks for mutual information estimation.
result MINE is linearly scalable in dimensionality and sample size, and is strongly consistent.
Paper mines rank data patterns from rankings.
problem Mining rank data patterns from rankings.
method Proposes algorithms for frequent rankings and dependencies.
result Experimental validation of algorithms on synthetic and real data.
Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft wares are available, but data mining in agricultural soil datasets is a relatively…
LetSIP learns relevant patterns for user interests in data mining.
problem Redundancy in pattern mining makes it hard for analysts to identify relevant patterns.
method Combines pattern sampling with interactive data mining, using user feedback to learn sampling distribution.
result Favourable trade-offs in quality-diversity and exploitation-exploration compared to existing methods.
This paper tracks coin circulation in Bitcoin to identify miners and analyze mining pool structures.
problem Identifying and understanding Bitcoin miners and their profit distribution schemes.
method Constructs fresh coin circulation networks and uses a heuristic algorithm to compare networks from different mining pools.
result Infers common profit distribution schemes of Bitcoin mining pools and observes an increasing trend in miner numbers.
Real-world data typically contain repeated and periodic patterns. This suggests that they can be effectively represented and compressed using only a few coefficients of an appropriate basis (e.g., Fourier, Wavelets, etc.). However, distance estimation when the data are represented using different sets of coefficients i…
Study shows Bitcoin mining with surplus electricity can boost KEPCO's financial stability.
problem Improving energy resource efficiency and reducing KEPCO's debt.
method Utilized surplus electricity for Bitcoin mining using Antminer S21 XP Hyd, analyzed with Random Forest Regressor and Long Short-Term Memory models.
result Bitcoin mining with surplus electricity generates economic revenue, minimizes energy loss, and resolves payment issues for KEPCO.
This work improves deep metric learning by online soft mining and class-aware attention.
problem Slow convergence and poor performance due to a large fraction of trivial samples in deep metric learning.
method Online Soft Mining (OSM) and Class-Aware Attention (CAA) to select and focus on relevant samples.
result Significantly outperforms state-of-the-art methods on fine-grained visual categorization and video-based person re-identification datasets.
Distributed, online data mining systems have emerged as a result of applications requiring analysis of large amounts of correlated and high-dimensional data produced by multiple distributed data sources. We propose a distributed online data classification framework where data is gathered by distributed data sources and…
A new framework for mining high utility patterns in interval-based sequences.
problem Mining patterns in events that persist over varying time intervals and considering event utility.
method Integrates utility into interval-based sequences and proposes HUIPMiner algorithm with pruning strategy.
result HUIPMiner efficiently finds high utility patterns in real datasets.