Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,786 papers · 148 categories

Trend · papers per month

5.0%10.1%15.1%20.1% · Jun 202319922001200920172026
48 results for correlated random variables

LMMVAE improves VAE for correlated data by separating latent variables into fixed and random parts.

problem Correlated data in tabular and image datasets.
method Integrates random effects into VAE architecture, separating latent variables into fixed and random parts.
result Significant improvement in reconstruction error and likelihood loss on unseen data.

A new method estimates conditional canonical correlations using random forests.

problem Estimating relationships between two sets of variables given covariates.
method Random Forest with Canonical Correlation Analysis (RFCCA)
result RFCCA provides accurate canonical correlation estimations and well-controlled Type-1 error.

The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.

problem The standard Pearson correlation coefficient is limited to two variables and doesn't meet the needs for multi-variable analysis.
method The authors use random matrix theory to extend Pearson's correlation coefficient to an arbitrary number of variables.
result The extended correlation coefficient is useful for gauging noise and selecting features, particularly in classification.

Two new methods for analyzing repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.

problem Analyzing complex data structures with multiple features over time.
method Two generalizations of canonical correlation analysis for repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
result Consistency rates for transformation and correlation estimators, relaxing common assumptions.

FREEtree improves tree-based methods for correlated longitudinal data.

problem Poor performance of Random Forests in high dimensional longitudinal data with correlated features.
method FREEtree uses a piecewise random effects model and clustering with WGCNA to select features and maintain interpretability.
result FREEtree outperforms other tree-based methods in prediction and feature selection accuracy.

LMLFM tackles predictive modeling from longitudinal data with mixed correlations.

problem Learning predictive models from longitudinal data with complex correlations and non-linear interactions.
method Longitudinal Multi-Level Factorization Machine (LMLFM) that selects predictive fixed and random effects.
result LMLFM outperforms state-of-the-art methods in predictive accuracy, variable selection, and scalability.

Better signal detection in undersampled data using joint and cross covariances.

problem Detecting shared signals in high-dimensional data with limited samples.
method Analysis of three covariance matrices: individual, cross, and joint.
result Joint and cross covariance matrices detect signals earlier than individual covariances.

We present a general method to detect and extract from a finite time sample statistically meaningful correlations between input and output variables of large dimensionality. Our central result is derived from the theory of free random matrices, and gives an explicit expression for the interval where singular values are…

2005-12-10abs ↗pdf ↗

We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-Rényi Maximum Correlation Coefficient. RDC is defined in terms of correlation of random non-linear copula projections; it is invariant with respec…

2013-04-29abs ↗pdf ↗

Non-symmetric rectangular correlation matrices occur in many problems in economics. We test the method of extracting statistically meaningful correlations between input and output variables of large dimensionality and build a toy model for artificially included correlations in large random time series.The results are t…

2010-04-26abs ↗pdf ↗

While records and order statistics of independent and identically distributed (i.i.d.) random variables X_1, ..., X_N are fully understood, much less is known for strongly correlated random variables, which is often the situation encountered in statistical physics. Recently, it was shown, in a series of works, that one…

2013-05-03abs ↗pdf ↗

Lognormal random variables appear naturally in many engineering disciplines, including wireless communications, reliability theory, and finance. So, too, does the sum of (correlated) lognormal random variables. Unfortunately, no closed form probability distribution exists for such a sum, and it requires approximation. …

2015-08-30abs ↗pdf ↗

This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single trea…

2011-02-06abs ↗pdf ↗

In this paper, a class of statistics named ART (the alternant recursive topology statistics) is proposed to measure the properties of correlation between two variables. A wide range of bi-variable correlations both linear and nonlinear can be evaluated by ART efficiently and equitably even if nothing is known about the…

2016-01-07abs ↗pdf ↗

The paper develops a test for independence of selected Gaussian variables after thresholding correlations.

problem Testing independence of selected Gaussian variables after thresholding correlations.
method The approach involves conditioning on the selection event and using a new characterization of the conditioning event in terms of canonical correlation.
result The proposed test has higher power than a naive approach that ignores selection effects.

Proposes a new method to estimate variable importance in black box models, mitigating correlation effects.

problem Correlation between covariates affects the interpretation of variable importance parameters.
method Develops a modified LOCO (Leave Out COvariates) method and uses semiparametric models for estimation.
result Shows how to estimate a modified LOCO method that mitigates correlation effects.

The problem of using observed correlations to infer causal relations is relevant to a wide variety of scientific disciplines. Yet given correlations between just two classical variables, it is impossible to determine whether they arose from a causal influence of one on the other or a common cause influencing both, unle…

2014-06-19abs ↗pdf ↗

Reduces selection bias in estimating individual treatment effects.

problem Selection bias in counterfactual reasoning.
method Auto-encoder with regularized loss based on Pearson Correlation Coefficient.
result Improves performance in estimating individual treatment effects.

We investigate financial market correlations using random matrix theory and principal component analysis. We use random matrix theory to demonstrate that correlation matrices of asset price changes contain structure that is incompatible with uncorrelated random price changes. We then identify the principal components o…

2010-11-14abs ↗pdf ↗

NURD improves model performance by distilling representations independent of nuisance variables.

problem Models trained under spurious correlations may fail on data with different nuisance-label relationships.
method Developed Nuisance-Randomized Distillation (NURD) to find representations independent of nuisance variables.
result NURD finds representations that perform better regardless of nuisance-label relationships.

We present sharp tail asymptotics for the density and the distribution function of linear combinations of correlated log-normal random variables, that is, exponentials of components of a correlated Gaussian vector. The asymptotic behavior turns out to depend on the correlation between the components, and the explicit s…

2013-09-12abs ↗pdf ↗

The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.

problem Uncertainty in predicting macroeconomic variables like prices and returns.
method Defines theoretical lower bounds of uncertainty and upper limits on forecast accuracy based on statistical moments and trade volumes.
result Accuracy of forecasts of probabilities of macroeconomic variables doesn't exceed Gaussian approximations.

We introduce Network Maximal Correlation (NMC) as a multivariate measure of nonlinear association among random variables. NMC is defined via an optimization that infers transformations of variables by maximizing aggregate inner products between transformed variables. For finite discrete and jointly Gaussian random vari…

2016-06-15abs ↗pdf ↗

Gaussian copulas are widely used in the industry to correlate two random variables when there is no prior knowledge about the co-dependence between them. The perturbed Gaussian copula approach allows introducing the skew information of both random variables into the co-dependence structure. The analytical expression of…

2010-02-27abs ↗pdf ↗

Simple conditions for comonotonic additive risk measures from acceptance sets.

problem Conditions for comonotonic additive risk measures from acceptance sets.
method Conditions on acceptance sets for induced comonotonic additive risk measures.
result Acceptance sets induce comonotonic additive risk measures if and only if the acceptance sets and their complements are stable under convex combinations of comonotonic random variables.

This paper compares imputation and direct parameter estimation methods for missing data in correlation matrix visualization.

problem Missing data challenges in estimating correlation coefficients for accurate visualization.
method Comparison of imputation and direct parameter estimation methods for handling missing data.
result Direct parameter estimation (DPER) outperforms imputation for accurate correlation matrix visualization.

This paper presents Correlated Nystrom Views (XNV), a fast semi-supervised algorithm for regression and classification. The algorithm draws on two main ideas. First, it generates two views consisting of computationally inexpensive random features. Second, XNV applies multiview regression using Canonical Correlation Ana…

2013-06-24abs ↗pdf ↗

In this manuscript, we analytically and numerically study statistical properties of an heteroskedastic process based on the celebrated ARCH generator of random variables whose variance is defined by a memory of qmq_{m}-exponencial, form (eqm=1x=exe_{q_{m}=1}^{x}=e^{x}). Specifically, we inspect the self-correlation function o…

2008-06-16abs ↗pdf ↗

Neural networks learn faster with correlated latent variables.

problem Efficiently learning from higher-order correlations in neural networks.
method Analytical derivation and simulations of two-layer neural networks.
result Correlations between latent variables speed up learning from higher-order correlations.

Study measures uncertainty in MST identification across different correlation networks.

problem Uncertainty in MST identification across various correlation-based market networks.
method Developed a framework using random variable networks (RVN) to measure uncertainty of MST identification.
result FDR is the most appropriate measure for MST identification reliability.

Inference and learning of graphical models are both well-studied problems in statistics and machine learning that have found many applications in science and engineering. However, exact inference is intractable in general graphical models, which suggests the problem of seeking the best approximation to a collection of …

2015-02-03abs ↗pdf ↗

Researchers identify critical protein residues using advanced graph theory.

problem Identifying essential residues in proteins for function.
method Learning Random Geometric Graphs (RGG) with Cramer's V correlation and organic thresholding.
result Advanced RGG methods accurately identify critical residues compared to existing techniques.

This paper presents a new ensemble learning method for classification problems called projection pursuit random forest (PPF). PPF uses the PPtree algorithm introduced in Lee et al. (2013). In PPF, trees are constructed by splitting on linear combinations of randomly chosen variables. Projection pursuit is used to choos…

2018-07-19abs ↗pdf ↗

The study improves the perceptron's storage capacity by optimizing variable selection.

problem Distinguishing genuine structure from random correlations in high-dimensional data.
method Replica method from statistical mechanics for optimal variable selection.
result Optimal variable selection can surpass the Cover--Gardner bound for pattern classification.

Inference and learning of graphical models are both well-studied problems in statistics and machine learning that have found many applications in science and engineering. However, exact inference is intractable in general graphical models, which suggests the problem of seeking the best approximation to a collection of …

2010-11-15abs ↗pdf ↗