The generalized correlation approach, which has been successfully used in statistical radio physics to describe non-Gaussian random processes, is proposed to describe stochastic financial processes. The generalized correlation approach has been used to describe a non-Gaussian random walk with independent, identically d…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A detailed analysis of correlation between stock returns at high frequency is compared with simple models of random walks. We focus in particular on the dependence of correlations on time scales - the so-called Epps effect. This provides a characterization of stochastic models of stock price returns which is appropriat…
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
The study finds significant power-law cross correlations in Bitcoin's return-volatility dynamics.
Simple model finds high correlation in retail crypto returns.
This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single trea…
This paper improves multi-label classification by leveraging high-order label correlations.
New model analyzes dynamic correlations in stock returns.
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
Model explains Netflix Prize success with structured correlation.
New estimator reveals intraday betas mainly driven by correlations.
DTCCA learns nonlinear transformations of multi-view data for high-order correlation.
In this paper we use wavelet concepts to show that correlation coefficient between two financial data's is not constant but varies with scale from high correlation value to strongly anti-correlation value This studies is important because correlation coefficient is used to quantify degree of independence between two va…
Canonical Correlation Analysis (CCA) is a classical tool for finding correlations among the components of two random vectors. In recent years, CCA has been widely applied to the analysis of genomic data, where it is common for researchers to perform multiple assays on a single set of patient samples. Recent work has pr…
We develop a framework for analyzing extreme values in correlated financial data.
We introduce a method to predict which correlation matrix coefficients are likely to change their signs in the future in the high-dimensional regime, i.e. when the number of features is larger than the number of samples per feature. The stability of correlation signs, two-by-two relationships, is found to depend on thr…
Proposes FarmHazard model for hazard regression with correlated covariates.
Spatially relaxed inference tackles high-dimensional linear models with correlated covariates.
Proposes a method to select features for deep learning in noisy, high-dimensional data.
New methods test correlation between network structure and node features.
The study identifies spurious correlations in high-dimensional regression and quantifies their impact.
We describe a new optimization scheme for finding high-quality correlation clusterings in planar graphs that uses weighted perfect matching as a subroutine. Our method provides lower-bounds on the energy of the optimal correlation clustering that are typically fast to compute and tight in practice. We demonstrate our a…
This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally dependent dimensions, compare the advantages of each test statistic, examine thei…
A new method for real-time CCA on streaming data.
We obtain general, exact formulas for the overlaps between the eigenvectors of large correlated random matrices, with additive or multiplicative noise. These results have potential applications in many different contexts, from quantum thermalisation to high dimensional statistics. We find that the overlaps only depend …
New framework selects high-quality pretraining data without training LLMs.
New insights into ridge regression with correlated data, improving risk prediction.
SGE-Kriging reduces high-dimensional surrogate modelling costs.
In a very high-dimensional vector space, two randomly-chosen vectors are almost orthogonal with high probability. Starting from this observation, we develop a statistical factor model, the random factor model, in which factors are chosen at random based on the random projection method. Randomness of factors has the con…
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
Canonical correlation analysis was proposed by Hotelling [6] and it measures linear relationship between two multidimensional variables. In high dimensional setting, the classical canonical correlation analysis breaks down. We propose a sparse canonical correlation analysis by adding l1 constraints on the canonical vec…
We introduce a method to learn a hierarchy of successively more abstract representations of complex data based on optimizing an information-theoretic objective. Intuitively, the optimization searches for a set of latent factors that best explain the correlations in the data as measured by multivariate mutual informatio…
Cluster GARCH model improves multivariate GARCH for high-dimensional asset returns.
We propose improved methods to identify stock groups using the correlation matrix of stock price changes. By filtering out the marketwide effect and the random noise, we construct the correlation matrix of stock groups in which nontrivial high correlations between stocks are found. Using the filtered correlation matrix…
Study finds multifractal cross-correlations between agricultural markets and external uncertainties.
A new screening rule improves lasso solving speed.
Paper evaluates and improves private feature selection methods.
New hierarchical model improves on standard practice for high-dimensional data.
It is commonly believed that the correlations between stock returns increase in high volatility periods. We investigate how much of these correlations can be explained within a simple non-Gaussian one-factor description with time independent correlations. Using surrogate data with the true market return as the dominant…
We propose a correlated stochastic process of which the novel non-Gaussian probability mass function is constructed by exactly solving moment generating function. The calculation of cumulants and auto-correlation shows that the process is convergent and scale invariant in the large but finite number limit. We demonstra…
Sparse GCA finds linear relationships in multiple datasets, using gradient descent.
We analyze the daily stock data of the Nasdaq Composite index in the 22-year period 1992-2013 and identify market states as clusters of correlation matrices with similar correlation structures. We investigate the stability of the correlation structure of each state by estimating the statistical fluctuations of correlat…
Improved CEM for fast real-time planning in high-dimensional control tasks.
We derive high-order compact finite difference schemes for option pricing in stochastic volatility models on non-uniform grids. The schemes are fourth-order accurate in space and second-order accurate in time for vanishing correlation. In our numerical study we obtain high-order numerical convergence also for non-zero …
Neurons in the visual cortex are correlated in their variability. The presence of correlation impacts cortical processing because noise cannot be averaged out over many neurons. In an effort to understand the functional purpose of correlated variability, we implement and evaluate correlated noise models in deep convolu…
A new screening method for high-dimensional data reduces computational cost.
Algorithm recovers permutations of high-dimensional Gaussian vectors with constant correlation.
It has been shown that instead of learning actual object features, deep networks tend to exploit non-robust (spurious) discriminative features that are shared between training and test sets. Therefore, while they achieve state of the art performance on such test sets, they achieve poor generalization on out of distribu…