Pearson distance fails to be a metric, but can be fixed.
problem The Pearson distance is not a metric, leading to issues in applications.
method Showed that Pearson distance is not a metric and provided alternative metrics.
result Pearson distance can be repaired by using alternative metrics derived from the correlation coefficient.
For time series comparisons, it has often been observed that z-score normalized Euclidean distances far outperform the unnormalized variant. In this paper we show that a z-score normalized, squared Euclidean Distance is, in fact, equal to a distance based on Pearson Correlation. This has profound impact on many distanc…
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
Robust test for distributions under Hellinger distance, simpler than optimal tests.
problem Testing and estimating distributions robustly under Hellinger distance.
method Simple robust hypothesis test with optimal sample complexity, robust to Hellinger distance perturbations.
result Empirically demonstrated robustness and power of the test on canonical distributions.
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
New bounds for Neyman-Pearson region using f-divergences.
problem Bounding the Neyman-Pearson region for hypothesis testing.
method Establishing novel lower and upper bounds using f-divergences. result Best possible lower bound for the Neyman-Pearson boundary using hockey-stick f-divergences. The study uses DCC for financial market analysis, revealing hidden correlations.
problem Identifying hidden nonlinear correlations in financial markets.
method Agglomerative hierarchical clustering with distance correlation coefficient.
result DCC reveals more information than Pearson correlation for financial data.
A graph traversal algorithm for cold-start news recommendation using named entities.
problem Cold-start news recommendation for articles without user-specific information.
method Graph traversal algorithm and novel weighting scheme for named entities over a knowledge graph.
result Our method produces stronger Pearson correlation to human similarity scores than other cold-start methods.
The article generalizes Pearson correlation to Riemannian manifolds.
problem Analyzing statistical models on non-linear manifolds.
method Reconstitutes Pearson correlation properties and derives a nonlinear generalization.
result Developed the Riemann-Pearson Correlation for manifold analysis.
Paper proposes Gini distance statistics for estimating feature-label dependence.
problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.
This paper presents a new approach for filter design based on stochastic distances and tests between distributions. A window is defined around each pixel, overlapping samples are compared and only those which pass a goodness-of-fit test are used to compute the filtered value. The technique is applied to intensity SAR d…
We compute exact values respectively bounds of "distances" - in the sense of (transforms of) power divergences and relative entropy - between two discrete-time Galton-Watson branching processes with immigration GWI for which the offspring as well as the immigration is arbitrarily Poisson-distributed (leading to arbitra…
Financial markets analyzed by reducing correlation matrix complexity.
problem Understanding complex financial market correlations.
method Coarse graining Pearson correlation matrices into Guhr matrices by market sectors.
result Significant reduction in the number of relevant variables.
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
problem The standard Pearson correlation coefficient is limited to two variables and doesn't meet the needs for multi-variable analysis.
method The authors use random matrix theory to extend Pearson's correlation coefficient to an arbitrary number of variables.
result The extended correlation coefficient is useful for gauging noise and selecting features, particularly in classification.
Robust hypothesis testing designs a test for worst-case distributions using kernel methods.
problem Design a robust test for hypothesis testing under uncertainty sets.
method Data-driven uncertainty sets constructed using kernel mean embeddings and maximum mean discrepancy (MMD). Bayesian and Neyman-Pearson settings investigated.
result Proposed robust kernel tests are exponentially consistent and asymptotically optimal.
Model predicts epileptic seizures with high accuracy using EEG signals.
problem Predicting epileptic seizures with high accuracy for diagnosis and treatment.
method Pearson's product-moment correlation coefficient with a linear classifier on generalized Gaussian modeling.
result 100% effectiveness for sensitivity and specificity greater than 83%.
Adapts Neyman-Pearson classification for both source and target distribution shifts.
problem Minimizing errors while controlling both Type-I and Type-II errors under distribution shifts.
method Derives an adaptive procedure that guarantees improved error rates and adapts to uninformative sources.
result Automatic adaptation to uninformative sources avoids negative transfer.
We examine the efficiency of the Asymmetric Power ARCH (APARCH) model in the case where the residuals follow the standardized Pearson type IV distribution. The model is tested with a variety of loss functions and the efficiency is examined via application of several statistical tests and risk measures. The results indi…
This paper presents a new approach for filter design based on stochastic distances and tests between distributions. A window is defined around each pixel, samples are compared and only those which pass a goodness-of-fit test are used to compute the filtered value. The technique is applied to intensity Synthetic Apertur…
This work shows cosine similarity is equivalent to Pearson correlation for word vectors, but not all vectors are suitable for cosine.
problem The use of cosine similarity for semantic textual similarity is often taken for granted, despite its limitations.
method Characterized cases where Pearson correlation is unfit and introduced rank correlation as an alternative.
result Pearson correlation is equivalent to cosine similarity for many word vectors but not all, and rank correlation can improve performance.
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
problem Asymmetric binary classification problems with unequal error severities.
method Develops TUBE-CS algorithm to bridge cost-sensitive and Neyman-Pearson paradigms.
result High-probability control of population type I error.
RS-Del provides robustness for sequence classifiers against edit distance attacks.
problem Certifying robustness of discrete sequence classifiers against edit distance attacks.
method Randomized deletion (RS-Del) for discrete sequence classifiers, focusing on edit distance-bounded adversaries.
result Achieved a certified accuracy of 91% at an edit distance radius of 128 bytes on malware detection.
Neyman-Pearson testing improves goodness of fit in detecting new physics.
problem Detecting small anomalies in data distributions.
method Employing Neyman-Pearson strategy with a rich parametrized family of models.
result Neyman-Pearson testing is more sensitive to small departures and unbiased towards specific anomalies.
Unified framework for various probability distribution distances.
problem Handling diverse probability distribution distances in statistics.
method General framework covering density-based and distribution-function-based divergences.
result Unified approach to classical and modern statistical procedures.
Entropy measures in their various incarnations play an important role in the study of stochastic time series providing important insights into both the correlative and the causative structure of the stochastic relationships between the individual components of a system. Recent applications of entropic techniques and th…
Characterizes distribution-free rates in unbalanced classification problems.
problem Minimizing error under two different distributions in unbalanced settings.
method Characterizes minimax rates over all pairs of distributions using a geometric condition.
result Identifies a dichotomy between hard and easy classes based on a three-points-separation condition.
This study uses local Gaussian correlation to analyze stock return tails, revealing more sensitive network properties.
problem Misleading results from Pearson correlation in financial networks.
method Local Gaussian correlation coefficient for capturing nonlinear dependence and heavy-tailed distributions.
result Local Gaussian correlation network among negative tails is more sensitive to stock market risks.
In this work, a novel solution to the speaker identification problem is proposed through minimization of statistical divergences between the probability distribution (g). of feature vectors from the test utterance and the probability distributions of the feature vector corresponding to the speaker classes. This approac…
Novel method prices call options using Pearson diffusion processes.
problem Pricing European call options with skewness and kurtosis.
method Modeling asset returns with Pearson diffusion processes.
result Proposed method outperforms Black-Scholes and Heston models.
The paper tackles Neyman-Pearson classification control issues.
problem Neyman-Pearson classification's control constraint is hard to satisfy in finite samples.
method Developed refined learning procedures under two accuracy control strategies.
result Proposed methods achieve desired control levels in finite samples.
The paper uses distance correlation for brain connectivity and a novel multi-task learning model for age prediction.
problem Estimating age-related gender differences in brain functional connectivity.
method Estimates functional connectivity using distance correlation and proposes a non-convex multi-task learning model.
result The proposed non-convex multi-task learning model outperforms other models in age prediction and gender-specific connectivity.
This paper presents two approaches for filter design based on stochastic distances for intensity speckle reduction. A window is defined around each pixel, overlapping samples are compared and only those which pass a goodness-of-fit test are used to compute the filtered value. The tests stem from stochastic divergences …
USP test improves on Pearson's chi-squared and G-test for independence.
problem Deficiencies in Pearson's chi-squared and G-test for independence. method USP test based on U-statistic estimator of population dependence measure. result USP test controls size, handles small cell counts, and detects minimal violations of independence.
In this short report, we investigate the ability of the DCCA coefficient to measure correlation level between non-stationary series. Based on a wide Monte Carlo simulation study, we show that the DCCA coefficient can estimate the correlation coefficient accurately regardless the strength of non-stationarity (measured b…
Unified framework for Bayes-optimal classifiers under group fairness.
problem Mitigating disparate impacts from algorithmic predictions in high-stakes decision-making.
method Unified framework based on Neyman-Pearson argument for deriving Bayes-optimal classifiers under group fairness constraints.
result Proposes FairBayes method that directly controls disparity and achieves optimal fairness-accuracy tradeoff.
New method corrects bias in density ratio estimation for missing data.
problem Missing data bias in density ratio estimation.
method Adapted KLIEP method (M-KLIEP) for MNAR data.
result M-KLIEP restores consistency and minimax optimality.
Develops NPMC method for noisy labels, improving multiclass classification accuracy.
problem Asymmetric misclassification costs and label noise in multiclass classification.
method Empirical likelihood approach using exponential tilting density ratio model.
result Root n consistent and asymptotically normal estimators for clean labels and noise mechanism.
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
problem Asymmetric misclassification costs in multi-class classification problems.
method Establishes connection with cost-sensitive learning, proposes two algorithms, extends NP oracle properties.
result Proposes algorithms with theoretical guarantees for multi-class Neyman-Pearson classification.
The paper develops approximations for Pearson's chi-square statistic and applies them to confidence intervals.
problem Finding confidence intervals for strictly convex functions of discrete distribution weights.
method Non-asymptotic local normal approximation for multinomial probabilities, deriving bounds and coupling inequalities.
result Developed methods to find confidence intervals for negative entropy of discrete distributions.
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
Enhanced metrics for multiclass classification improve on existing methods.
problem Lack of decisive poor classification results in existing multiclass metrics.
method Introduces three new metrics derived from multivariate Pearson correlation coefficients.
result New metrics decisively indicate poor classification results.
A new method detects and displays pairwise dependence between variates.
problem Detecting and visualizing dependence between variates of different types.
method Recursive random binning with approximations to Pearson's statistic.
result The method is well-calibrated and powerful against common test alternatives.
Stock price movement reveals complex interdependencies that are simplified through linear correlation.
problem Exploring the spectral dynamics of the Indonesian capital market using structural network representations.
method Combining three dependency estimators (Pearson, MI adaptive binning, and MI-kNN) with two graph filtering schemes (MST and PMFG) and four community decoders.
result MI adaptive binning is shown to be more proportional than kNN for detecting residual information.
A neural network for online NP classification with reduced complexity.
problem Online nonlinear Neyman-Pearson classification.
method Single hidden layer feedforward neural network (SLFN) initialized with random Fourier features (RFFs). Uses stochastic gradient descent for sequential learning.
result Expedited online adaptation and powerful nonlinear Neyman-Pearson modeling.
This paper uses rank correlation methods to construct MSTs from financial returns, finding them more stable and robust.
problem Stability and robustness of MSTs constructed from financial correlation matrices.
method Pearson, Spearman, and Kendall's τ rank correlation methods applied to daily financial returns. result Rank MSTs are more stable and robust than MSTs constructed using Pearson correlation.
The paper studies statistical properties of CART regression trees.
problem Understanding the statistical properties of CART regression trees.
method The paper constructs a prior distribution on split points and solves a nonlinear optimization problem to bound the Pearson correlation between the optimal decision stump and response data.
result CART with cost-complexity pruning achieves an optimal complexity/goodness-of-fit tradeoff when the depth scales with the logarithm of the sample size.
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …