The article generalizes Pearson correlation to Riemannian manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Adapts Neyman-Pearson classification for both source and target distribution shifts.
The paper develops approximations for Pearson's chi-square statistic and applies them to confidence intervals.
USP test improves on Pearson's chi-squared and -test for independence.
We examine the efficiency of the Asymmetric Power ARCH (APARCH) model in the case where the residuals follow the standardized Pearson type IV distribution. The model is tested with a variety of loss functions and the efficiency is examined via application of several statistical tests and risk measures. The results indi…
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
The paper tackles Neyman-Pearson classification control issues.
A new method detects and displays pairwise dependence between variates.
Neyman-Pearson testing improves goodness of fit in detecting new physics.
Investigation of the market graph attracts a growing attention in market network analysis. One of the important problem connected with market graph is to identify it from observations. Traditional way for the market graph identification is to use a simple procedure based on statistical estimations of Pearson correlatio…
The paper studies statistical properties of CART regression trees.
This study uses local Gaussian correlation to analyze stock return tails, revealing more sensitive network properties.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
New bounds for Neyman-Pearson region using -divergences.
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
The gain-loss asymmetry, observed in the inverse statistics of stock indices is present for logarithmic return levels that are over , and it is the result of the non-Pearson type auto-correlations in the index. These non-Pearson type correlations can be viewed also as functionally dependent daily volatilities, ext…
This paper examines autocorrelation in major crypto markets, finding persistent correlations on short time frames.
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
Optimal selective classification using likelihood ratios improves model reliability.
Recently the interest of researchers has shifted from the analysis of synchronous relationships of financial instruments to the analysis of more meaningful asynchronous relationships. Both of those analyses are concentrated only on Pearson's correlation coefficient and thus intraday lead-lag relationships associated wi…
Financial markets analyzed by reducing correlation matrix complexity.
The paper analyzes skewness and kurtosis measures for skew-elliptical distributions.
Model predicts epileptic seizures with high accuracy using EEG signals.
Correlation matrices play a key role in many multivariate methods (e.g., graphical model estimation and factor analysis). The current state-of-the-art in estimating large correlation matrices focuses on the use of Pearson's sample correlation matrix. Although Pearson's sample correlation matrix enjoys various good prop…
For time series comparisons, it has often been observed that z-score normalized Euclidean distances far outperform the unnormalized variant. In this paper we show that a z-score normalized, squared Euclidean Distance is, in fact, equal to a distance based on Pearson Correlation. This has profound impact on many distanc…
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
Paper introduces an online method for estimating the difference between two probability distributions.
A new method reduces feature screening cost from to .
This paper considers an often forgotten relationship, the time delay between a cause and its effect in economies and finance. We treat the case of Foreign Direct Investment (FDI) and economic growth, - measured through a country Gross Domestic Product (GDP). The pertinent data refers to 43 countries, over 1970-2015, - …
A new method validates generative models in high-dimensional data.
New RDPC dissimilarity measure improves time series clustering.
Entropy measures in their various incarnations play an important role in the study of stochastic time series providing important insights into both the correlative and the causative structure of the stochastic relationships between the individual components of a system. Recent applications of entropic techniques and th…
Characterizes distribution-free rates in unbalanced classification problems.
The Pearson distance between a pair of random variables with correlation , namely, 1-, has gained widespread use, particularly for clustering, in areas such as gene expression analysis, brain imaging and cyber security. In all these applications it is implicitly assumed/required that the distance …
Paper optimizes statistical estimation for randomized smoothing to reduce adversarial robustness certification time.
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
Novel method prices call options using Pearson diffusion processes.
The objective of change-point detection is to discover abrupt property changes lying behind time-series data. In this paper, we present a novel statistical change-point detection algorithm based on non-parametric divergence estimation between time-series samples from two retrospective segments. Our method uses the rela…
In this short report, we investigate the ability of the DCCA coefficient to measure correlation level between non-stationary series. Based on a wide Monte Carlo simulation study, we show that the DCCA coefficient can estimate the correlation coefficient accurately regardless the strength of non-stationarity (measured b…
Unified framework for Bayes-optimal classifiers under group fairness.
New method corrects bias in density ratio estimation for missing data.
Unified platform for statistical and machine learning in bioinformatics.
Develops NPMC method for noisy labels, improving multiclass classification accuracy.
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
Empirical evidence is given for a significant difference in the collective trend of the share prices during the stock index rising and falling periods. Data on the Dow Jones Industrial Average and its stock components are studied between 1991 and 2008. Pearson-type correlations are computed between the stocks and avera…
The statistical analysis of discrete data has been the subject of extensive statistical research dating back to the work of Pearson. In this survey we review some recently developed methods for testing hypotheses about high-dimensional multinomials. Traditional tests like the test and the likelihood ratio test ca…
Enhanced metrics for multiclass classification improve on existing methods.
A large body of research into semantic textual similarity has focused on constructing state-of-the-art embeddings using sophisticated modelling, careful choice of learning signals and many clever tricks. By contrast, little attention has been devoted to similarity measures between these embeddings, with cosine similari…