In the last few years, many different performance measures have been introduced to overcome the weakness of the most natural metric, the Accuracy. Among them, Matthews Correlation Coefficient has recently gained popularity among researchers not only in machine learning but also in several application fields such as bio…
Enhanced metrics for multiclass classification improve on existing methods.
problem Lack of decisive poor classification results in existing multiclass metrics.
method Introduces three new metrics derived from multivariate Pearson correlation coefficients.
result New metrics decisively indicate poor classification results.
Improved multiclass classification with class-weighted nearest neighbors.
problem Multiclass classification with large or imbalanced classes.
method Class-weighted k-nearest neighbors algorithm, derived bounds on accuracy and risk.
result Optimized classification metrics like F1 score or Matthew's Correlation Coefficient.
Logit-link models reveal socio-temporal effects on microfinance delinquency.
problem Understanding and quantifying socio-temporal factors affecting microfinance loan delinquency.
method Developed and evaluated discrete-time logit-link models with fixed-effects and frailty extensions.
result Simple random intercept structures capture latent heterogeneity in microfinance repayment behavior.
Deep neural networks improve heart disease diagnosis accuracy.
problem Improving accuracy of heart disease diagnosis.
method Design and use of deep neural networks (DNNs) for detecting heart disease based on clinical data.
result HEARO-5 architecture yields 99% accuracy and 0.98 MCC.
Novel algorithm optimizes decision trees for nonlinear metrics.
problem Optimizing decision trees for nonlinear metrics like F1-score.
method Bi-objective optimisation approach to find optimal trees on Pareto frontier.
result The optimal tree for nonlinear metrics lies on the Pareto frontier.
This paper visualizes uncertainty in classifier performance metrics.
problem Overemphasis on model performance metrics risks overlooking uncertainty.
method Developed visualizations of confusion matrix metric distributions.
result Uncertainty in performance metrics can overshadow model differences.
Predict stock price movements using financial data and news articles with LLMs.
problem Predicting stock price movements using financial data and news articles.
method Combining financial data and news articles, employing pre-trained LLMs, and using retrieval augmentation techniques.
result Predicted stock price movements with a weighted F1-score of 58.5% and 59.1%.
The paper discusses thresholds and bounds for accuracy in binary classification systems.
problem The accuracy of binary classification systems and its dependence on prevalence.
method Analyzing the precision-prevalence curve and negative predictive value-prevalence curve to find thresholds and bounds.
result Thresholds (φe and φn) bound various accuracy metrics (Fβ, F1, FM, MCC) and the ratio of maximum accuracy to prevalence. Study improves machine learning models for GI tract disease detection using comprehensive evaluations and cross-dataset testing.
problem Incomplete or incorrect evaluation of machine learning models for GI tract diseases.
method Comprehensive evaluations of five machine learning models using Global Features and Deep Neural Networks, introducing performance hexagons and cross-dataset testing.
result Demonstrates the need for more sophisticated performance metrics and evaluation methods to build generalizable models.
New metrics improve performance in imbalanced classification problems.
problem Established metrics favor classifiers ignoring minority classes.
method Introduce robust modifications of F-score and MCC.
result TPR is bounded away from 0 in imbalanced settings.
Research predicts healthcare index movements using historical OHLC data.
problem Predicting the directional movement of healthcare indices based on historical data.
method Supervised classification task with a one-step-ahead rolling window, using a diverse feature set including OHLC ratios.
result Robust predictive performance with accuracy exceeding 0.8 and Matthews correlation coefficients above 0.6, highlighting the importance of nowcasting features.
Forest tree species mapped with high accuracy using satellite data.
problem Classifying dominant tree species in Swedish forests.
method Extreme gradient boosting model with Bayesian optimization, combining Sentinel-1/2 satellite data and field observations.
result Overall accuracy of 85%, F1 score of 0.82, Matthews correlation coefficient of 0.81.
Bayesian network models are finding success in characterizing enzyme-catalyzed reactions, slow conformational changes, predicting enzyme inhibition, and genomics. In this work, we apply them to statistical modeling of peptides by simultaneously identifying amino acid sequence motifs and using a motif-based model to cla…
Linear classifiers separate the data with a hyperplane. In this paper we focus on the novel method of construction of multithreshold linear classifier, which separates the data with multiple parallel hyperplanes. Proposed model is based on the information theory concepts -- namely Renyi's quadratic entropy and Cauchy-S…
Neural network improves breast cancer diagnosis with high accuracy.
problem Improving accuracy in breast cancer diagnosis.
method Higher-order probabilistic perceptron (HOPP) model.
result HOPP model achieves up to 97% accuracy in classifying breast cancer tumors.
The paper evaluates six imbalanced data strategies across 58 datasets.
problem Mitigating imbalanced data in binary classification problems.
method Comprehensive comparative analysis of 10 under-sampling, 5 over-sampling, 2 ensemble, and 3 specialized algorithms on 58 real-life datasets.
result The effectiveness of imbalanced data strategies varies significantly depending on the performance metric used.
In this paper we use wavelet concepts to show that correlation coefficient between two financial data's is not constant but varies with scale from high correlation value to strongly anti-correlation value This studies is important because correlation coefficient is used to quantify degree of independence between two va…
Deep autoencoder predicts cancer types from DNA methylation patterns.
problem Differentiating cancer types based on DNA methylation states.
method Deep learning system with CpG island state classification and statistical methods.
result Overall Sensitivity of 88.24%, Specificity of 83.33%, Accuracy of 84.75%.
In this short report, we investigate the ability of the DCCA coefficient to measure correlation level between non-stationary series. Based on a wide Monte Carlo simulation study, we show that the DCCA coefficient can estimate the correlation coefficient accurately regardless the strength of non-stationarity (measured b…
A new weighted MCC measure improves classifier performance evaluation.
problem Lack of measures sensitive to observation weights in multiclass classification.
method Proposes weighted versions of Pearson-Matthews Correlation Coefficient (MCC) for binary and multiclass classification.
result Weighted MCC values are higher for classifiers that perform better on highly weighted observations.
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
Improved portfolio optimization using Kendall-like correlation coefficients.
problem Accurate estimation of eigenvectors in data-poor regimes for portfolio optimization.
method Developed generalized correlation coefficients based on Kendall's rank correlation.
result Markowitz portfolios with lower out-of-sample risk using these coefficients.
The study uses DCC for financial market analysis, revealing hidden correlations.
problem Identifying hidden nonlinear correlations in financial markets.
method Agglomerative hierarchical clustering with distance correlation coefficient.
result DCC reveals more information than Pearson correlation for financial data.
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
problem The standard Pearson correlation coefficient is limited to two variables and doesn't meet the needs for multi-variable analysis.
method The authors use random matrix theory to extend Pearson's correlation coefficient to an arbitrary number of variables.
result The extended correlation coefficient is useful for gauging noise and selecting features, particularly in classification.
Machine learning predicts trauma patient mortality risk.
problem Predicting mortality risk in trauma patients using traditional regression models.
method Transfer learning-based machine learning algorithm applied to trauma patient data.
result Machine learning model achieved similar performance to contemporary models without restrictive criteria.
Advanced AI model predicts stock movements post earnings reports.
problem Inaccurate stock predictions from earnings reports.
method Fine-tuned LLMs with QLoRA compression, integrating financial and market data.
result Significantly improved predictive accuracy compared to benchmarks.
Standardizes weighted ranking correlation coefficients to maintain zero expected value.
problem Measuring correlation between weighted rankings of items.
method Develops a standardization function g(·) that transforms coefficients to zero expected value under randomness.
result A general standardization function g(Γ) that preserves the domain [-1,1] and reduces to the identity for coefficients already satisfying zero-expected-value property.
In the paper, we introduce a new measure of correlation between possibly non-stationary series. As the measure is based on the detrending moving-average cross-correlation analysis (DMCA), we label it as the DMCA coefficient ρDMCA(λ) with a moving average window length λ. We analytically show that the coefficient…
Method predicts which high-dimensional correlation signs will change in the future.
problem Predicting which correlation matrix coefficients will change signs in high-dimensional data.
method Stability of correlation signs depends on three-by-three relationships, inspired by Heider social cohesion theory.
result The method accurately predicts the stability of correlation signs in high-dimensional data.
New correlation measures improve classifier performance assessment.
problem Improving assessment of classifiers and raters.
method Introducing CO-, ANTI-, and COANTI-correlation coefficients.
result Demonstrated new measures are powerful for classifying confusion matrices.
Objective: In this work, we perform margin assessment of human breast tissue from optical coherence tomography (OCT) images using deep neural networks (DNNs). This work simulates an intraoperative setting for breast cancer lumpectomy. Methods: To train the DNNs, we use both the state-of-the-art methods (Weight Decay an…
When common factors strongly influence two power-law cross-correlated time series recorded in complex natural or social systems, using classic detrended cross-correlation analysis (DCCA) without considering these common factors will bias the results. We use detrended partial cross-correlation analysis (DPXA) to uncover…
Survey of recent measures of association, including a new coefficient.
problem Exploring new measures of association in statistics.
method Survey and introduction of a new correlation coefficient.
result Proposed a new extension of the correlation coefficient to standard Borel spaces.
Study examines NFT market dynamics using correlation and noise analysis.
problem Understanding correlations and noise in NFT market.
method Used detrended correlation coefficient and correlation matrix analysis.
result Correlation strength in NFT market is lower than in cryptocurrency markets.
We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-Rényi Maximum Correlation Coefficient. RDC is defined in terms of correlation of random non-linear copula projections; it is invariant with respec…
The detrended cross-correlation coefficient ρDCCA has recently been proposed to quantify the strength of cross-correlations on different temporal scales in bivariate, non-stationary time series. It is based on the detrended cross-correlation and detrended fluctuation analyses (DCCA and DFA, respectively) and c…
Graph-based ML improves defect prediction in software development.
problem Challenges in predicting defect-prone changes in complex software development.
method Building contribution graphs from developers and source files, using graph-based ML for defect prediction.
result Graph-based ML leads to significantly better defect prediction (F1 score up to 77.55%, MCC up to 53.16%).
We consider the effects of the global financial crisis through a local Korean financial market around the 2008 crisis. We analyze 185 individual stock prices belonging to the KOSPI (Korea Composite Stock Price Index), cosidering three time periods: the time before, during, and after the crisis. The complex networks gen…
Benchmark evaluates AI-generated financial QA hallucinations, highlighting system vulnerabilities.
problem Ensuring factual accuracy of AI-generated financial QA outputs.
method Developed a benchmark dataset and evaluated six detection methods under clean and noisy conditions.
result LLM-based judges and embedding methods perform best, but degrade under noisy conditions.
Discovering a correlation from one variable to another variable is of fundamental scientific and practical interest. While existing correlation measures are suitable for discovering average correlation, they fail to discover hidden or potential correlations. To bridge this gap, (i) we postulate a set of natural axioms …
This work shows cosine similarity is equivalent to Pearson correlation for word vectors, but not all vectors are suitable for cosine.
problem The use of cosine similarity for semantic textual similarity is often taken for granted, despite its limitations.
method Characterized cases where Pearson correlation is unfit and introduced rank correlation as an alternative.
result Pearson correlation is equivalent to cosine similarity for many word vectors but not all, and rank correlation can improve performance.
Paper presents a technique using Spearman's Rank Correlation Coefficient for KE in TDs.
problem Extracting common characteristics and grouping similar TDs.
method Spearman's Rank Correlation Coefficient (SRCC) for KE.
result SRCC proves a comprehensive measure for high-quality KE.
Predicting the price correlation of two assets for future time periods is important in portfolio optimization. We apply LSTM recurrent neural networks (RNN) in predicting the stock price correlation coefficient of two individual stocks. RNNs are competent in understanding temporal dependencies. The use of LSTM cells fu…
EAST aligns neural network classifiers with user-defined evaluation metrics.
problem Mismatch between neural network training and evaluation metrics leads to suboptimal performance.
method EAST uses dynamic thresholding, soft-set confusion matrix, and annealing to align neural network predictions with target evaluation metrics.
result EAST improves alignment between training objectives and evaluation metrics, outperforming existing methods.
Study uses machine learning to predict potato clones suitable for processing.
problem Efficiently identifying high-yield, disease-resistant potato varieties.
method Leveraged machine learning algorithms on Russet potato clones data from Oregon.
result Non-linear models like SVM and HGBC outperform traditional linear models in agricultural trials.
This study uses local Gaussian correlation to analyze stock return tails, revealing more sensitive network properties.
problem Misleading results from Pearson correlation in financial networks.
method Local Gaussian correlation coefficient for capturing nonlinear dependence and heavy-tailed distributions.
result Local Gaussian correlation network among negative tails is more sensitive to stock market risks.
Model predicts epileptic seizures with high accuracy using EEG signals.
problem Predicting epileptic seizures with high accuracy for diagnosis and treatment.
method Pearson's product-moment correlation coefficient with a linear classifier on generalized Gaussian modeling.
result 100% effectiveness for sensitivity and specificity greater than 83%.