We prove a central limit theorem for the components of the largest eigenvectors of the adjacency matrix of a finite-dimensional random dot product graph whose true latent positions are unknown. In particular, we follow the methodology outlined in \citet{sussman2012universally} to construct consistent estimates for the …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Novel risk matrix for optimal portfolio choice with tail risk considerations.
A new method for streaming PCA provides confidence intervals for eigenvector entries.
Spectral clustering is widely used to partition graphs into distinct modules or communities. Existing methods for spectral clustering use the eigenvalues and eigenvectors of the graph Laplacian, an operator that is closely associated with random walks on graphs. We propose a new spectral partitioning method that exploi…
New algorithm updates eigenvectors of evolving graphs efficiently.
Proposes a new algorithm to estimate invariant subspaces across multilayer networks.
We prove a central limit theorem for the components of the eigenvectors corresponding to the largest eigenvalues of the normalized Laplacian matrix of a finite dimensional random dot product graph. As a corollary, we show that for stochastic blockmodel graphs, the rows of the spectral embedding of the normalized La…
We quantify uncertainty in Oja's algorithm's leading eigenvector estimation.
Complex network analysis reveals dominant stocks in financial stock returns correlations.
New method detects global factors near BBP phase transition in high-dimensional data.
Improved portfolio optimization using Kendall-like correlation coefficients.
As relational datasets modeled as graphs keep increasing in size and their data-acquisition is permeated by uncertainty, graph-based analysis techniques can become computationally and conceptually challenging. In particular, node centrality measures rely on the assumption that the graph is perfectly known -- a premise …
A new algorithm reduces data dimensionality and decorrelation in a distributed setting.
In an era where accumulating data is easy and storing it inexpensive, feature selection plays a central role in helping to reduce the high-dimensionality of huge amounts of otherwise meaningless data. In this paper, we propose a graph-based method for feature selection that ranks features by identifying the most import…
Many pattern recognition methods rely on statistical information from centered data, with the eigenanalysis of an empirical central moment, such as the covariance matrix in principal component analysis (PCA), as well as partial least squares regression, canonical-correlation analysis and Fisher discriminant analysis. R…
New sampling methods improve node embedding efficiency.
The interbank market is considered one of the most important channels of contagion. Its network representation, where banks and claims/obligations are represented by nodes and links (respectively), has received a lot of attention in the recent theoretical and empirical literature, for assessing systemic risk and identi…
Graph embedding method captures both local and global network structure.
Network metrics form a fundamental part of the network analysis toolbox. Used to quantitatively measure different aspects of the network, these metrics can give insights into the underlying network structure and function. In this work, we connect network metrics to modern probabilistic machine learning. We focus on the…
Study eigenvector overlaps in large Gaussian matrices, simplifying for GOE.
Machine learning models perform better with location coordinates alone, not Moran Eigenvectors.
Paper addresses eigenvector perturbation in small eigen-gap scenarios.
We calculate eigenvector overlaps between intersecting time periods of covariance matrices.
In many applications, one has side information, e.g., labels that are provided in a semi-supervised manner, about a specific target region of a large data set, and one wants to perform machine learning and data analysis tasks "nearby" that prespecified target region. For example, one might be interested in the clusteri…
In spectral clustering, one defines a similarity matrix for a collection of data points, transforms the matrix to get the Laplacian matrix, finds the eigenvectors of the Laplacian matrix, and obtains a partition of the data using the leading eigenvectors. The last step is sometimes referred to as rounding, where one ne…
We study the problem asking if one can embed manifolds into finite dimensional Euclidean spaces by taking finite number of eigenvector fields of the connection Laplacian. This problem is essential for the dimension reduction problem in massive data analysis. Singer-Wu proposed the vector diffusion map which embeds mani…
New metric tensor field on symmetric matrices simplifies eigenvector computation.
Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.
New neural architectures invariant to sign flips and basis symmetries for graph representation learning.
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
We demonstrate the existence of an empirical linkage between the nominal financial networks and the underlying economic fundamentals across countries. We construct the nominal return correlation networks from daily data to encapsulate sector-level dynamics and figure the relative importance of the sectors in the nomina…
Estimating the leading principal components of data, assuming they are sparse, is a central task in modern high-dimensional statistics. Many algorithms were developed for this sparse PCA problem, from simple diagonal thresholding to sophisticated semidefinite programming (SDP) methods. A key theoretical question is und…
New theory for eigenvectors of generalized Laplacian matrices, addressing dependency issues.
Fast algorithm recovers principal eigenvector from noisy matrices.
This paper develops the exact linear relationship between the leading eigenvector of the unnormalized modularity matrix and the eigenvectors of the adjacency matrix. We propose a method for approximating the leading eigenvector of the modularity matrix, and we derive the error of the approximation. There is also a comp…
New method improves subspace iteration for eigenvectors in machine learning.
The paper explores how kernel eigenalignments affect generalization in KRR.
The problem of estimating sparse eigenvectors of a symmetric matrix attracts a lot of attention in many applications, especially those with high dimensional data set. While classical eigenvectors can be obtained as the solution of a maximization problem, existing approaches formulate this problem by adding a penalty te…
Improved spectral clustering with fewer eigenvectors performs better.
New insights into spectral clustering reveal strong connections within eigenvectors.
The original contributions of this paper are twofold: a new understanding of the influence of noise on the eigenvectors of the graph Laplacian of a set of image patches, and an algorithm to estimate a denoised set of patches from a noisy image. The algorithm relies on the following two observations: (1) the low-index e…
We characterize the contractions that are similar to the backward shift in the Hardy space . This characterization is given in terms of the geometry of the eigenvector bundles of the operators.
Graph convolutional networks fail to use eigenvectors beyond the first, unlike spectral embedding.
We provide new examples of diffusion operators in dimension 2 and 3 which have orthogonal polynomials as eigenvectors. Their construction rely on the finite subgroups of O(3) and their invariant polynomials.
Study eigenvalues and eigenvectors in neural networks, focusing on signal propagation.
The paper tackles learning symmetries in data without expert knowledge.
This paper aims to address two fundamental challenges arising in eigenvector estimation and inference for a low-rank matrix from noisy observations: (1) how to estimate an unknown eigenvector when the eigen-gap (i.e. the spacing between the associated eigenvalue and the rest of the spectrum) is particularly small; (2) …
New method improves covariance estimation for weighted samples.