Paper presents a copula-based method to efficiently generate correlated sample paths from multi-step time series models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Enhances labels from unlabeled data using sample correlations.
We describe a method to determine the eigenvalue density of empirical covariance matrix in the presence of correlations between samples. This is a straightforward generalization of the method developed earlier by the authors for uncorrelated samples. The method allows for exact determination of the experimental spectru…
Paper addresses the disparity between sampled and mean representations in disentangled learning.
Unsupervised anomaly detection aims to identify anomalous samples from highly complex and unstructured data, which is pervasive in both fundamental research and industrial applications. However, most existing methods neglect the complex correlation among data samples, which is important for capturing normal patterns fr…
Improved sample complexity for Gaussian Mixture Models using Pair Correlation Factor.
Estimates covariance matrices with correlations between samples.
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
Correlation matrices play a key role in many multivariate methods (e.g., graphical model estimation and factor analysis). The current state-of-the-art in estimating large correlation matrices focuses on the use of Pearson's sample correlation matrix. Although Pearson's sample correlation matrix enjoys various good prop…
The Epps effect varies under different sampling schemes, affecting correlation emergence rates.
This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single trea…
We study the problem of finding the most mutually correlated arms among many arms. We show that adaptive arms sampling strategies can have significant advantages over the non-adaptive uniform sampling strategy. Our proposed algorithms rely on a novel correlation estimator. The use of this accurate estimator allows us t…
Enhances sensitivity analysis for correlated inputs.
Machine learning often needs to model density from a multidimensional data sample, including correlations between coordinates. Additionally, we often have missing data case: that data points can miss values for some of coordinates. This article adapts rapid parametric density estimation approach for this purpose: model…
We study finite sample properties of estimators of power-law cross-correlations -- detrended cross-correlation analysis (DCCA), height cross-correlation analysis (HXA) and detrending moving-average cross-correlation analysis (DMCA) -- with a special focus on short-term memory bias as well as power-law coherency. Presen…
Portfolio allocation and risk management make use of correlation matrices and heavily rely on the choice of a proper correlation matrix to be used. In this regard, one important question is related to the choice of the proper sample period to be used to estimate a stable correlation matrix. This paper addresses this qu…
New insights into ridge regression with correlated data, improving risk prediction.
Paper develops efficient algorithms for learning rationalizable equilibria in multiplayer games.
Correlation matrices are omnipresent in multivariate data analysis. When the number d of variables is large, the sample estimates of correlation matrices are typically noisy and conceal underlying dependence patterns. We consider the case when the variables can be grouped into K clusters with exchangeable dependence; t…
Logit correction improves model performance by correcting spurious correlations.
Paper proposes PSIPS for identifying Pareto set with correlated objectives.
We study Principal Component Analysis (PCA) in a setting where a part of the corrupting noise is data-dependent and, as a result, the noise and the true data are correlated. Under a bounded-ness assumption on the true data and the noise, and a simple assumption on data-noise correlation, we obtain a nearly optimal samp…
Mean representations of VAEs are correlated but still useful for tasks.
Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. In this paper, we establish a new framework that generalizes distance correlation --- a correlation mea…
Study sharpens threshold for matching correlated graphs without labels.
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
Mitigates spurious correlations without bias labels.
Slice Sampling has emerged as a powerful Markov Chain Monte Carlo algorithm that adapts to the characteristics of the target distribution with minimal hand-tuning. However, Slice Sampling's performance is highly sensitive to the user-specified initial length scale hyperparameter and the method generally struggles with …
Given two data matrices and , sparse canonical correlation analysis (SCCA) is to seek two sparse canonical vectors and to maximize the correlation between and . However, classical and sparse CCA models consider the contribution of all the samples of data matrices and thus cannot identify an unde…
We study the sample complexity of canonical correlation analysis (CCA), \ie, the number of samples needed to estimate the population canonical correlation and directions up to arbitrarily small error. With mild assumptions on the data distribution, we show that in order to achieve -suboptimality in a properly define…
Improved best-arm identification in correlated multi-armed bandits.
New SMC sampler improves diffusion model sampling efficiency.
Losaw improves FI scores by decorrelating features in ML models.
A new algorithm removes unexpected correlations in biased data for better clustering.
VADD enhances discrete diffusion models by capturing inter-dimensional correlations, improving sample quality.
We propose a novel approach for sampling realistic financial correlation matrices. This approach is based on generative adversarial networks. Experiments demonstrate that generative adversarial networks are able to recover most of the known stylized facts about empirical correlation matrices estimated on asset returns.…
Proposes meTS for efficient exploration in correlated bandits.
Diffusion models learn simple statistics before complex ones, revealing a sample complexity exponent.
We consider the problem of providing nonparametric confidence guarantees for undirected graphs under weak assumptions. In particular, we do not assume sparsity, incoherence or Normality. We allow the dimension to increase with the sample size . First, we prove lower bounds that show that if we want accurate infe…
This work explains how maximizing latent correlations across multiple data views helps in identifying shared and private components.
Discovering a correlation from one variable to another variable is of fundamental scientific and practical interest. While existing correlation measures are suitable for discovering average correlation, they fail to discover hidden or potential correlations. To bridge this gap, (i) we postulate a set of natural axioms …
Correlations and other collective phenomena in a schematic model of heterogeneous binary agents (individual spin-glass samples) are considered on the complete graph and also on 2d and 3d regular lattices. The system's stochastic dynamics is studied by numerical simulations. The dynamics is so slow that one can meaningf…
Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is…
Private method measures nonlinear correlations between data hosted across two entities.
Sparse GCA finds linear relationships in multiple datasets, using gradient descent.
Neurons in the visual cortex are correlated in their variability. The presence of correlation impacts cortical processing because noise cannot be averaged out over many neurons. In an effort to understand the functional purpose of correlated variability, we implement and evaluate correlated noise models in deep convolu…
Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.
PROBE optimizes best-arm identification with cheap proxies, improving sample complexity.