Fast online algorithm for nonparametric correlations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recent advances in topic models have explored complicated structured distributions to represent topic correlation. For example, the pachinko allocation model (PAM) captures arbitrary, nested, and possibly sparse correlations between topics using a directed acyclic graph (DAG). While PAM provides more flexibility and gr…
Proposes a new method to estimate variable importance in black box models, mitigating correlation effects.
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
We are often interested in explaining data through a set of hidden factors or features. When the number of hidden features is unknown, the Indian Buffet Process (IBP) is a nonparametric latent feature model that does not bound the number of active features in dataset. However, the IBP assumes that all latent features a…
We develop correlated random measures, random measures where the atom weights can exhibit a flexible pattern of dependence, and use them to develop powerful hierarchical Bayesian nonparametric models. Hierarchical Bayesian nonparametric models are usually built from completely random measures, a Poisson-process based c…
Canonical Correlation Analysis (CCA) is a classical tool for finding correlations among the components of two random vectors. In recent years, CCA has been widely applied to the analysis of genomic data, where it is common for researchers to perform multiple assays on a single set of patient samples. Recent work has pr…
Unified framework for interpreting complex regression models with many predictors.
Canonical correlation analysis (CCA) is a classical representation learning technique for finding correlated variables in multi-view data. Several nonlinear extensions of the original linear CCA have been proposed, including kernel and deep neural network methods. These approaches seek maximally correlated projections …
New estimator for joint entropy outperforms existing methods in various distributions.
The paper introduces a new method for modeling correlations in high-dimensional data.
CATVI improves variational inference for Bayesian nonparametric models by reducing divergence and improving prediction accuracy.
Method predicts which high-dimensional correlation signs will change in the future.
New method improves portfolio allocation using local Gaussian correlation.
In this paper, a class of statistics named ART (the alternant recursive topology statistics) is proposed to measure the properties of correlation between two variables. A wide range of bi-variable correlations both linear and nonlinear can be evaluated by ART efficiently and equitably even if nothing is known about the…
Model dynamic customer sensitivities across categories.
This paper solves nonparametric estimation of continuous DPPs using kernel methods.
In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the rankings of the values of two variables from a data set with a large number n of o…
Genetic sequence data are well described by hidden Markov models (HMMs) in which latent states correspond to clusters of similar mutation patterns. Theory from statistical genetics suggests that these HMMs are nonhomogeneous (their transition probabilities vary along the chromosome) and have large support for self tran…
GNet uses Gaussian processes for scalable, flexible neural networks.
GNet uses Gaussian processes for scalable, flexible neural networks.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
Multilayer bootstrap network builds a gradually narrowed multilayer nonlinear network from bottom up for unsupervised nonlinear dimensionality reduction. Each layer of the network is a nonparametric density estimator. It consists of a group of k-centroids clusterings. Each clustering randomly selects data points with r…
Proposes a method for valid inference in GPLSIMs with longitudinal data.
New Hermite series estimator for Spearman rank correlation in non-stationary data.
A new nonparametric test measures dependence between variables using decision trees.
Estimates self- and cross-impact concavity and decay patterns in financial markets.
Parametric divergences are effective for generative modeling despite being non-optimal.
Sum-Product Networks simplify building models for mixed data types.
Proposes logistic-beta process for modeling dependent probabilities with beta marginals.
Method estimates observation functions in state-space models without supervision.
The paper presents a novel approach to multi-output regression using probabilistic circuits.
The paper proposes a method to construct well-calibrated prediction sets for correlated target variables.
A two-step nonparametric method estimates financial systemic risk.
We consider the problem of providing nonparametric confidence guarantees for undirected graphs under weak assumptions. In particular, we do not assume sparsity, incoherence or Normality. We allow the dimension to increase with the sample size . First, we prove lower bounds that show that if we want accurate infe…
MT-HAL learns features and task associations for multiple tasks with a shared sparse structure.
Geometric QHD tests improve hub detection in correlated data.
This appendix proves CORN's universal consistency. One of Bin's PhD thesis examiner (Special thanks to Vladimir Vovk from Royal Holloway, University of London) suggested that CORN is universal and provided sketch proof of Lemma 1.6, which is the key of this proof. Based on the proof in Gyprfi et al. [2006], we thus pro…
This paper deals with the problem of nonparametric independence testing, a fundamental decision-theoretic problem that asks if two arbitrary (possibly multivariate) random variables are independent or not, a question that comes up in many fields like causality and neuroscience. While quantities like correlation o…
One popular approach to model the limit order books dynamics of the best bid and ask at level-1 is to use the reduced-form diffusion approximations. It is well known that the biggest contributing factor to the price movement is the imbalance of the best bid and ask. We investigate the data of the level-1 limit order bo…
The article compares predictor importance in classification problems with categorical outcomes.
Genome-wide association studies have proven to be essential for understanding the genetic basis of disease. However, many complex traits---personality traits, facial features, disease subtyping---are inherently high-dimensional, impeding simple approaches to association mapping. We developed a nonparametric Bayesian re…
Nonnegative Matrix Factorization (NMF) aims to factorize a matrix into two optimized nonnegative matrices appropriate for the intended applications. The method has been widely used for unsupervised learning tasks, including recommender systems (rating matrix of users by items) and document clustering (weighting matrix …
We present the discrete infinite logistic normal distribution (DILN), a Bayesian nonparametric prior for mixed membership models. DILN is a generalization of the hierarchical Dirichlet process (HDP) that models correlation structure between the weights of the atoms at the group level. We derive a representation of DILN…
Unsupervised two-view learning, or detection of dependencies between two paired data sets, is typically done by some variant of canonical correlation analysis (CCA). CCA searches for a linear projection for each view, such that the correlations between the projections are maximized. The solution is invariant to any lin…
Method learns software resource usage from snapshots.
We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal family proposed by Liu et al (2009). In terms of estimation, we exploit nonparametric rank-based correlation coefficient estimators includin…
A new method for approximating softmax and Gaussian kernels with reduced error.