The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The study uses DCC for financial market analysis, revealing hidden correlations.
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. In this paper, we establish a new framework that generalizes distance correlation --- a correlation mea…
For time series comparisons, it has often been observed that z-score normalized Euclidean distances far outperform the unnormalized variant. In this paper we show that a z-score normalized, squared Euclidean Distance is, in fact, equal to a distance based on Pearson Correlation. This has profound impact on many distanc…
This note improves correlation stress tests using geodesic distance.
Paper relaxes differential privacy for correlated features, improving privacy-utility trade-off.
The problem of filtering information from large correlation matrices is of great importance in many applications. We have recently proposed the use of the Kullback-Leibler distance to measure the performance of filtering algorithms in recovering the underlying correlation matrix when the variables are described by a mu…
We show that the Kullback-Leibler distance is a good measure of the statistical uncertainty of correlation matrices estimated by using a finite set of data. For correlation matrices of multivariate Gaussian variables we analytically determine the expected values of the Kullback-Leibler distance of a sample correlation …
The paper uses distance correlation for brain connectivity and a novel multi-task learning model for age prediction.
Reduces data leakage in distributed deep learning models.
Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is…
Proposes a hierarchical clustering method for positive and negative dissimilarities.
Private method measures nonlinear correlations between data hosted across two entities.
In this paper, we deal with the problem of inferring causal directions when the data is on discrete domain. By considering the distribution of the cause and the conditional distribution mapping cause to effect as independent random variables, we propose to infer the causal direction via comparing the di…
BDC uses Distance Correlation for efficient Bayesian optimization of expensive functions.
Develops a method for stress testing correlations of financial portfolios.
Geography effect is investigated for the Chinese stock market including the Shanghai and Shenzhen stock markets, based on the daily data of individual stocks. The Shanghai city and the Guangdong province can be identified in the stock geographical sector. By investigating a geographical correlation on a geographical pa…
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
In our work, we propose a novel formulation for supervised dimensionality reduction based on a nonlinear dependency criterion called Statistical Distance Correlation, Szekely et. al. (2007). We propose an objective which is free of distributional assumptions on regression variables and regression model assumptions. Our…
In this on-going work, I explore certain theoretical and empirical implications of data transformations under the PCA. In particular, I state and prove three theorems about PCA, which I paraphrase as follows: 1). PCA without discarding eigenvector rows is injective, but looses this injectivity when eigenvector rows are…
Modified cosine distance improves similarity performance in data with variance and correlation.
Study predicts climate data at distant locations using machine learning.
This paper analyzes correlations in patterns of trading of different members of the London Stock Exchange. The collection of strategies associated with a member institution is defined by the sequence of signs of net volume traded by that institution in hour intervals. Using several methods we show that there are signif…
Enhances community detection in correlated networks with node attributes.
Feature interactions can contribute to a large proportion of variation in many prediction models. In the era of big data, the coexistence of high dimensionality in both responses and covariates poses unprecedented challenges in identifying important interactions. In this paper, we suggest a two-stage interaction identi…
New unsupervised feature selection method for imbalanced datasets.
Paper uses news data to model asset correlations without market data.
DC-SIS selects features faster than mRMR for Parkinson's vocal diagnosis.
Testing two potentially multivariate variables for statistical dependence on the basis finite samples is a fundamental statistical challenge. Here we explore a family of tests that adapt to the complexity of the relationship between the variables, promising robust power across scenarios. Building on the distance correl…
New method disentangles correlated factors without independence assumption.
The medoid of a set of n points is the point in the set that minimizes the sum of distances to other points. It can be determined exactly in O(n^2) time by computing the distances between all pairs of points. Previous works show that one can significantly reduce the number of distance computations needed by adaptively …
The cluster analysis methods are used in order to perform a comparative study of 15 EU countries in relation with the fluctuations of some basic macroeconomic indicators. The statistical distances between countries are calculated for various moving time windows, and the time variation of the mean statistical distance i…
We discuss some methods to quantitatively investigate the properties of correlation matrices. Correlation matrices play an important role in portfolio optimization and in several other quantitative descriptions of asset price dynamics in financial markets. Specifically, we discuss how to define and obtain hierarchical …
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
A new algorithm of the analysis of correlation among economy time series is proposed. The algorithm is based on the power law classification scheme (PLCS) followed by the analysis of the network on the percolation threshold (NPT). The algorithm was applied to the analysis of correlations among GDP per capita time serie…
Geometric QHD tests improve hub detection in correlated data.
We have recently introduced the ``thermal optimal path'' (TOP) method to investigate the real-time lead-lag structure between two time series. The TOP method consists in searching for a robust noise-averaged optimal path of the distance matrix along which the two time series have the greatest similarity. Here, we gener…
Neurons in the visual cortex are correlated in their variability. The presence of correlation impacts cortical processing because noise cannot be averaged out over many neurons. In an effort to understand the functional purpose of correlated variability, we implement and evaluate correlated noise models in deep convolu…
This paper explores the relationships between migration and trade using a complex-network approach. We show that: (i) both weighted and binary versions of the networks of international migration and trade are strongly correlated; (ii) such correlations can be mostly explained by country economic/demographic size and ge…
In 2012, JPMorgan accumulated a USD~6.2 billion loss on a credit derivatives portfolio, the so-called `London Whale', partly as a consequence of de-correlations of non-perfectly correlated positions that were supposed to hedge each other. Motivated by this case, we devise a factor model for correlations that allows for…
s-OTDD compares datasets efficiently without training, robust to class variations.
Financial markets are well known examples of multi-fractal complex systems that have garnered much interest in their characterization through complex network theory. The recent studies have used correlation based distance metrics for defining and analyzing financial networks. In this work the singularity strength is em…
One major hurdle in the road toward a low carbon economy is the present entanglement of developed economies with oil. This tight relationship is mirrored in the correlation between most of economic indicators with oil price. This paper addresses the role of oil compared to the other three main energy commodities -coal,…
Improves joint distribution learning for high-dimensional datasets with complex correlations.
Reshef & Reshef recently published a paper in which they present a method called the Maximal Information Coefficient (MIC) that can detect all forms of statistical dependence between pairs of variables as sample size goes to infinity. While this method has been praised by some, it has also been criticized for its lack …
The Pearson distance between a pair of random variables with correlation , namely, 1-, has gained widespread use, particularly for clustering, in areas such as gene expression analysis, brain imaging and cyber security. In all these applications it is implicitly assumed/required that the distance …
Unsupervised learning of disentangled representations involves uncovering of different factors of variations that contribute to the data generation process. Total correlation penalization has been a key component in recent methods towards disentanglement. However, Kullback-Leibler (KL) divergence-based total correlatio…