Proposes a new regularizer for semi-supervised learning on multilayer graphs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method for distributed PCA using matrix β-mean.
Rk-means clusters relational data without full matrix, speeding up clustering.
New method allows generating independent data matrices from summary statistics.
Sharp inequalities for matrix means with unknown variance.
The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…
We show that the objective function of conventional k-means clustering can be expressed as the Frobenius norm of the difference of a data matrix and a low rank approximation of that data matrix. In short, we show that k-means clustering is a matrix factorization problem. These notes are meant as a reference and intende…
DeepTMR reorders matrices without prior knowledge of structural patterns.
Paper develops a method to construct confidence regions for model parameters using batch means method.
Improved method for computing Fréchet means on SPD matrices.
New method for matrix completion using Kronecker product approximation.
Novel mean estimation method under user-level differential privacy reduces noise in continual mean estimates.
We propose a penalized likelihood method to fit the linear discriminant analysis model when the predictor is matrix valued. We simultaneously estimate the means and the precision matrix, which we assume has a Kronecker product decomposition. Our penalties encourage pairs of response category mean matrices to have equal…
Signed graphs encode positive (attractive) and negative (repulsive) relations between nodes. We extend spectral clustering to signed graphs via the one-parameter family of Signed Power Mean Laplacians, defined as the matrix power mean of normalized standard and signless Laplacians of positive and negative edges. We pro…
A clustering algorithm uses the left Gram matrix for high dimensional data.
This paper solves the convergence problem for estimating MGGD parameters with a convex formulation.
Sharp threshold found for Frechet mean of inhomogeneous graphs.
Binary data matrices can represent many types of data such as social networks, votes, or gene expression. In some cases, the analysis of binary matrices can be tackled with nonnegative matrix factorization (NMF), where the observed data matrix is approximated by the product of two smaller nonnegative matrices. In this …
Missing data estimation is an important challenge with high-dimensional data arranged in the form of a matrix. Typically this data matrix is transposable, meaning that either the rows, columns or both can be treated as features. To model transposable data, we present a modification of the matrix-variate normal, the mea…
Kernel clustering algorithm improved for large datasets using incomplete Cholesky factorization.
Estimates low-rank distributional matrices from incomplete samples.
Study of metrics on positive-definite matrices from power potential, linking to power means.
We generalize the recently discovered relationship between JT gravity and double-scaled random matrix theory to the case that the boundary theory may have time-reversal symmetry and may have fermions with or without supersymmetry. The matching between variants of JT gravity and matrix ensembles depends on the assumed s…
Unified treatment of eigenvalue processes using Riemannian geometry.
The problem of low rank matrix completion is considered in this paper. To exploit the underlying low-rank structure of the data matrix, we propose a hierarchical Gaussian prior model, where columns of the low-rank matrix are assumed to follow a Gaussian distribution with zero mean and a common precision matrix, and a W…
Estimation of the covariance matrix has attracted a lot of attention of the statistical research community over the years, partially due to important applications such as Principal Component Analysis. However, frequently used empirical covariance estimator (and its modifications) is very sensitive to outliers in the da…
Paper proposes a method to break symmetries in Bayesian matrix factorization.
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
We consider the problem of estimating a consensus community structure by combining information from multiple layers of a multi-layer network using methods based on the spectral clustering or a low-rank matrix factorization. As a general theme, these "intermediate fusion" methods involve obtaining a low column rank matr…
Corrected whitening restores orthogonality in high-dimensional spherical Gaussian mixtures.
Enhances ROM simulation for multivariate systems with exact Kollo skewness.
We use the explicit relation between genus filtrated -loop means of the Gaussian matrix model and terms of the genus expansion of the Kontsevich--Penner matrix model (KPMM), which is the generating function for volumes of discretized (open) moduli spaces (discrete volumes), to express Gaussian means…
In this paper we derive the optimal linear shrinkage estimator for the high-dimensional mean vector using random matrix theory. The results are obtained under the assumption that both the dimension and the sample size tend to infinity in such a way that . Under weak conditions imposed on…
In this paper, we provide a unified analysis of temporal difference learning algorithms with linear function approximators by exploiting their connections to Markov jump linear systems (MJLS). We tailor the MJLS theory developed in the control community to characterize the exact behaviors of the first and second order …
In this work, the possibility of clustering correlated random variables was examined, both because of their mutual similarity and because of their similarity to the principal components. The k-means algorithm and spectral algorithms were used for clustering. For spectral methods, the similarity matrix was both the matr…
A method to remove mean-shift noise from PCA using knockoffs.
Principal components analysis (PCA) is a well-known technique for approximating a tabular data set by a low rank matrix. Here, we extend the idea of PCA to handle arbitrary data sets consisting of numerical, Boolean, categorical, ordinal, and other data types. This framework encompasses many well known techniques in da…
Since Li and Yau obtained the gradient estimate for the heat equation, related estimates have been extensively studied. With additional curvature assumptions, matrix estimates that generalize such estimates have been discovered for various time-dependent settings, including the heat equation on a Kähler manifold, Ricci…
We compute an approximate Fréchet mean for sets of sparse graphs.
We analyse the matrix factorization problem. Given a noisy measurement of a product of two matrices, the problem is to estimate back the original matrices. It arises in many applications such as dictionary learning, blind matrix calibration, sparse principal component analysis, blind source separation, low rank matrix …
Nonnegative Boltzmann machines (NNBMs) are recurrent probabilistic neural network models that can describe multi-modal nonnegative data. NNBMs form rectified Gaussian distributions that appear in biological neural network models, positive matrix factorization, nonnegative matrix factorization, and so on. In this paper,…
Paper analyzes VI for location-scale families, proving robustness guarantees for mean and correlation recovery.
Real-world data such as digital images, MRI scans and electroencephalography signals are naturally represented as matrices with structural information. Most existing classifiers aim to capture these structures by regularizing the regression matrix to be low-rank or sparse. Some other methodologies introduce factorizati…
Investigates portfolio optimization with and without gearing constraints.
It has been proposed that complex populations, such as those that arise in genomics studies, may exhibit dependencies among observations as well as among variables. This gives rise to the challenging problem of analyzing unreplicated high-dimensional data with unknown mean and dependence structures. Matrix-variate appr…
Method selects number of communities in weighted networks.
Paper improves matrix-valued data classification using nonparametric LDA.
The paper computes an approximation to the sample Frechet mean of graph sets using spectral information.