Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

2.4%4.8%7.1%9.5% · May 199719922001200920182026
48 results for correlation mining

The paper tackles sample complexity in high-dimensional data, focusing on correlation mining.

problem Understanding reliable inference in variable-rich, sample-starved data.
method Develops a unified statistical framework to quantify sample complexity for various inferential tasks.
result Illustrates high-dimensional learning rates and sample complexity for correlation mining.

Data mining techniques on the biological analysis are spreading for most of the areas including the health care and medical information. We have applied the data mining techniques, such as KNN, SVM, MLP or decision trees over a unique dataset, which is collected from 16,380 analysis results for a year. Furthermore we h…

2014-01-25abs ↗pdf ↗

The paper explores how mining costs, rewards, and blockchain security are interconnected.

problem Understanding the interdependencies between mining costs, mining rewards, and blockchain security.
method Theoretical derivation and empirical analysis using daily crypto market data and autoregressive distributed lag approach.
result Cryptocurrency price and mining rewards are intrinsically linked to blockchain security outcomes.

AlphaEval evaluates alpha mining models efficiently and comprehensively.

problem Lack of systematic evaluation for alpha mining models.
method Unified, parallelizable evaluation framework assessing predictive power, stability, robustness, financial logic, and diversity.
result AlphaEval achieves evaluation consistency comparable to comprehensive backtesting, providing more comprehensive insights and higher efficiency.

Randomization helps verify if data mining results are due to inherent patterns.

problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.

This paper offers a new algebraic perspective of GCCA using subspace intersection.

problem Finding common variables across multiple feature representations.
method Subspace intersection approach based on a (bi-)linear generative model.
result GCCA is equivalent to subspace intersection, with conditions for identifiable common subspace.

This paper improves image super-resolution by integrating cross-scale non-local attention.

problem Improving image super-resolution by leveraging long-range and cross-scale feature correlations.
method Proposes a Cross-Scale Non-Local (CS-NL) attention module integrated into a recurrent neural network.
result Significantly improved performance on SISR benchmarks.

AlphaSAGE mines diverse alphas via GFlowNets, overcoming RL issues.

problem Reward sparsity, inadequate sequential representations, and single optimal mode issues in RL for alphas.
method Structure-aware encoder (RGCN), GFlowNets, dense reward structure.
result Empirically outperforms existing baselines in mining diverse alphas.

A new framework forecasts stock trends by mining shared information from concepts.

problem Forecasting stock trends using static concept information limits accuracy.
method Proposes a graph-based framework that mines concept-oriented shared information from both predefined and hidden concepts.
result Improves stock trend forecasting performance through dynamic concept relevance and hidden concept information.

Paper proposes a new REINFORCE algorithm for mining formulaic alpha factors with reduced variance.

problem Mining formulaic alpha factors with interpretability and robustness in volatile markets.
method Developed a novel REINFORCE algorithm with a dedicated baseline and reward shaping.
result Boosts correlation with returns by 3.83% and enhances excess returns compared to existing methods.

RiskMiner discovers formulaic alphas using MCTS for better performance.

problem Mining formulaic alphas without considering structural information and alpha correlations.
method Formulates alpha mining as an MDP and solves it with a risk-seeking MCTS.
result Our method outperforms state-of-the-art benchmarks and achieves the most profitable results.

This study proposes a deep learning framework using ResNeXt for efficient financial data mining.

problem Complex financial data with high dimensionality, nonlinearity, and task correlations.
method Introduces ResNeXt into multi-task learning framework for efficient feature extraction and task collaboration.
result Significantly improved performance in classification and regression tasks on S&P 500 data.

ProxiModel extracts high-quality news events from news corpora.

problem Mining high-quality structured event knowledge from noisy news data.
method ProxiModel uses a proximity-network to model event correlation within and across news corpora.
result ProxiModel efficiently and effectively extracts high-quality event descriptors and attributes.

Feature selection has attracted significant attention in data mining and machine learning in the past decades. Many existing feature selection methods eliminate redundancy by measuring pairwise inter-correlation of features, whereas the complementariness of features and higher inter-correlation among more than two feat…

2015-02-01abs ↗pdf ↗

Discover novel multivariate relationships in time series data.

problem Capturing novel relationships between time series in complex systems.
method Introducing multipoles as linear relationships among more than two time series, identifying them as cliques of negative correlations in a correlation network.
result Almost all multipoles can be efficiently found using a clique-enumeration approach.

FactorMiner discovers financial alpha factors with low redundancy.

problem Finding novel financial alpha factors in a vast search space.
method Modular Skill Architecture and Experience Memory to distill and guide exploration.
result FactorMiner constructs a diverse library of high-quality factors with competitive performance.

A new imputation model for clinical data captures both cross-sectional and temporal correlations.

problem Missing values in multivariable time series clinical data.
method Integrates Gaussian processes with mixture models and individualized mixing weights.
result The proposed model provides more accurate imputation than benchmarks on real-world and synthetic datasets.

A neural network method to measure correlation among multiple variables.

problem Measuring correlation among multiple variables.
method Designing a neural network function, calculating differences, optimizing parameters, and using POET algorithm.
result The method can improve neural network performance and address issues like overfitting.

This study uses deep learning to analyze stock market sentiment from financial forums.

problem Improving stock market prediction accuracy through emotional analysis.
method Crawling financial forum data, training Bert model on financial corpus, using MIC for comparison.
result BERT model's emotional analysis of financial texts correlates with stock market fluctuations.

This paper analyzes the profitability of selfish mining on blockchain, considering the risk of ruin.

problem The profitability of selfish mining on blockchain, considering the risk of ruin.
method Formulated a stochastic model and used tools from applied probability and analysis to determine expected profit.
result Explicit expressions for expected profit under different scenarios were derived, identifying conditions for selfish mining as a strategic advantage.

LetSIP learns relevant patterns for user interests in data mining.

problem Redundancy in pattern mining makes it hard for analysts to identify relevant patterns.
method Combines pattern sampling with interactive data mining, using user feedback to learn sampling distribution.
result Favourable trade-offs in quality-diversity and exploitation-exploration compared to existing methods.

This paper tracks coin circulation in Bitcoin to identify miners and analyze mining pool structures.

problem Identifying and understanding Bitcoin miners and their profit distribution schemes.
method Constructs fresh coin circulation networks and uses a heuristic algorithm to compare networks from different mining pools.
result Infers common profit distribution schemes of Bitcoin mining pools and observes an increasing trend in miner numbers.

Study shows Bitcoin mining with surplus electricity can boost KEPCO's financial stability.

problem Improving energy resource efficiency and reducing KEPCO's debt.
method Utilized surplus electricity for Bitcoin mining using Antminer S21 XP Hyd, analyzed with Random Forest Regressor and Long Short-Term Memory models.
result Bitcoin mining with surplus electricity generates economic revenue, minimizes energy loss, and resolves payment issues for KEPCO.

This work improves deep metric learning by online soft mining and class-aware attention.

problem Slow convergence and poor performance due to a large fraction of trivial samples in deep metric learning.
method Online Soft Mining (OSM) and Class-Aware Attention (CAA) to select and focus on relevant samples.
result Significantly outperforms state-of-the-art methods on fine-grained visual categorization and video-based person re-identification datasets.

Distributed, online data mining systems have emerged as a result of applications requiring analysis of large amounts of correlated and high-dimensional data produced by multiple distributed data sources. We propose a distributed online data classification framework where data is gathered by distributed data sources and…

2013-07-02abs ↗pdf ↗

A new framework for mining high utility patterns in interval-based sequences.

problem Mining patterns in events that persist over varying time intervals and considering event utility.
method Integrates utility into interval-based sequences and proposes HUIPMiner algorithm with pruning strategy.
result HUIPMiner efficiently finds high utility patterns in real datasets.