Proposes PSCCA for estimating correlations and canonical correlations in sparse count data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper compares imputation and direct parameter estimation methods for missing data in correlation matrix visualization.
New trees-based models handle correlated data better.
We construct and analyze symmetrized delay correlation matrices for empirical data sets for atmopheric and financial data to derive information about correlation between different entities of the time series over time. The information about correlations is obtained by comparing the results for the eigenvalue distributi…
Proposes a multi-view VAE for imputing missing data from correlated sources.
Private method measures nonlinear correlations between data hosted across two entities.
In this paper we use wavelet concepts to show that correlation coefficient between two financial data's is not constant but varies with scale from high correlation value to strongly anti-correlation value This studies is important because correlation coefficient is used to quantify degree of independence between two va…
We examine Deep Canonically Correlated LSTMs as a way to learn nonlinear transformations of variable length sequences and embed them into a correlated, fixed dimensional space. We use LSTMs to transform multi-view time-series data non-linearly while learning temporal relationships within the data. We then perform corre…
In many scientific tasks we are interested in discovering whether there exist any correlations in our data. This raises many questions, such as how to reliably and interpretably measure correlation between a multivariate set of attributes, how to do so without having to make assumptions on distribution of the data or t…
CATS adapts multivariate time series models by addressing correlation shift.
Study identifies and analyzes spurious correlations in data-driven models.
Improved portfolio optimization using Kendall-like correlation coefficients.
Financial markets are highly correlated systems that reveal both the inter-market dependencies and the correlations among their different components. Standard analyzing techniques include correlation coefficients for pairs of signals and correlation matrices for rich multivariate data. In the latter case one constructs…
A new algorithm removes unexpected correlations in biased data for better clustering.
Proposes ACCA for better alignment of multiple data perspectives.
New Hermite series estimator for Spearman rank correlation in non-stationary data.
Study shows disentanglement models learn correlations from data, impacting fairness.
We study Principal Component Analysis (PCA) in a setting where a part of the corrupting noise is data-dependent and, as a result, the noise and the true data are correlated. Under a bounded-ness assumption on the true data and the noise, and a simple assumption on data-noise correlation, we obtain a nearly optimal samp…
Most data is multi-dimensional. Discovering whether any subset of dimensions, or subspaces, of such data is significantly correlated is a core task in data mining. To do so, we require a measure that quantifies how correlated a subspace is. For practical use, such a measure should be universal in the sense that it capt…
For multiple multivariate data sets, we derive conditions under which Generalized Canonical Correlation Analysis (GCCA) improves classification performance of the projected datasets, compared to standard Canonical Correlation Analysis (CCA) using only two data sets. We illustrate our theoretical results with simulation…
CSTS benchmarks time series clustering by evaluating correlation structures.
Proposes a method to enhance multi-view learning by maximizing higher order correlations.
A possible data source for the estimation of asset correlations is default time series. This study investigates the systematic error that is made if the exposure pool underlying a default time series is assumed to be homogeneous when in reality it is not. We find that the asset correlation will always be underestimated…
We review the decomposition method of stock return cross-correlations, presented previously for studying the dependence of the correlation coefficient on the resolution of data (Epps effect). Through a toy model of random walk/Brownian motion and memoryless renewal process (i.e. Poisson point process) of observation ti…
Method preserves correlations in synthetic data.
LMMVAE improves VAE for correlated data by separating latent variables into fixed and random parts.
Unified framework for generating data by modeling causal and correlational dependencies.
New method detects and analyzes correlation in multiple network data.
New method disentangles latent subspaces under correlation shifts.
Variational Auto-Encoders (VAEs) are capable of learning latent representations for high dimensional data. However, due to the i.i.d. assumption, VAEs only optimize the singleton variational distributions and fail to account for the correlations between data points, which might be crucial for learning latent representa…
We analyze the spectral properties of correlation matrices between distinct statistical systems. Such matrices are intrinsically non symmetric, and lend themselves to extend the spectral analyses usually performed on standard Pearson correlation matrices to the realm of complex eigenvalues. We employ some recent random…
Nonparametric correlations such as Spearman's rank correlation and Kendall's tau correlation are widely applied in scientific and engineering fields. This paper investigates the problem of computing nonparametric correlations on the fly for streaming data. Standard batch algorithms are generally too slow to handle real…
New estimator reveals intraday betas mainly driven by correlations.
New method corrects correlation bias in feature importance.
Generative Adversarial Networks (GAN) have shown great promise in tasks like synthetic image generation, image inpainting, style transfer, and anomaly detection. However, generating discrete data is a challenge. This work presents an adversarial training based correlated discrete data (CDD) generation model. It also de…
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
We apply random matrix theory to compare correlation matrix estimators C obtained from emerging market data. The correlation matrices are constructed from 10 years of daily data for stocks listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. We test the spectral properties of C against ra…
A new method for real-time CCA on streaming data.
Correlated component analysis as proposed by Dmochowski et al. (2012) is a tool for investigating brain process similarity in the responses to multiple views of a given stimulus. Correlated components are identified under the assumption that the involved spatial networks are identical. Here we propose a hierarchical pr…
Memory capacity of DAM scales exponentially with feature separation, unaffected by correlations.
Unsupervised anomaly detection aims to identify anomalous samples from highly complex and unstructured data, which is pervasive in both fundamental research and industrial applications. However, most existing methods neglect the complex correlation among data samples, which is important for capturing normal patterns fr…
Improved eigenvalue distribution method for financial data.
We propose a correlated stochastic process of which the novel non-Gaussian probability mass function is constructed by exactly solving moment generating function. The calculation of cumulants and auto-correlation shows that the process is convergent and scale invariant in the large but finite number limit. We demonstra…
A new method estimates conditional canonical correlations using random forests.
Two new methods for analyzing repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
The statistical dependencies which independent component analysis (ICA) cannot remove often provide rich information beyond the linear independent components. It would thus be very useful to estimate the dependency structure from data. While such models have been proposed, they usually concentrated on higher-order corr…
FREEtree improves tree-based methods for correlated longitudinal data.
TCGPN improves stock forecasting by capturing temporal correlation patterns.