Weight normalization speeds up matrix sensing problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
New method normalizes matrix features for robust low-rank approximation.
This paper develops the exact linear relationship between the leading eigenvector of the unnormalized modularity matrix and the eigenvectors of the adjacency matrix. We propose a method for approximating the leading eigenvector of the modularity matrix, and we derive the error of the approximation. There is also a comp…
In this paper we form relations for the determination of the elements of the Eötvös matrix of the Earth's normal gravity field. In addition a relation between the Gauss curvature of the normal equipotential surface and the Gauss curvature of the actual equipotential surface both passing through the point P is presented…
Matrix profile has been recently proposed as a promising technique to the problem of all-pairs-similarity search on time series. Efficient algorithms have been proposed for computing it, e.g., STAMP, STOMP and SCRIMP++. All these algorithms use the z-normalized Euclidean distance to measure the distance between subsequ…
We prove a central limit theorem for the components of the eigenvectors corresponding to the largest eigenvalues of the normalized Laplacian matrix of a finite dimensional random dot product graph. As a corollary, we show that for stochastic blockmodel graphs, the rows of the spectral embedding of the normalized La…
This work interprets diffusion score matching using normalizing flows for better model training and evaluations.
Doubly-stochastic normalization improves robustness to heteroskedastic noise.
NMF with specific constraints is equivalent to LDA.
New method clusters matrix-variate data with outliers.
New method for hyperparameter tuning in sparse matrix factorization.
New method proves asymptotic normality for matrix sensing problems.
Algorithm counts intersections of normal curves efficiently.
The paper studies matrix normalization and graph balancing using a new functional and gradient descent.
In recent years, data have become increasingly higher dimensional and, therefore, an increased need has arisen for dimension reduction techniques for clustering. Although such techniques are firmly established in the literature for multivariate data, there is a relative paucity in the area of matrix variate, or three-w…
Spectral clustering is a technique that clusters elements using the top few eigenvectors of their (possibly normalized) similarity matrix. The quality of spectral clustering is closely tied to the convergence properties of these principal eigenvectors. This rate of convergence has been shown to be identical for both th…
PSMM method optimizes matrix sufficient dimension reduction.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
The ubiquitous proliferation of online social networks has led to the widescale emergence of relational graphs expressing unique patterns in link formation and descriptive user node features. Matrix Factorization and Completion have become popular methods for Link Prediction due to the low rank nature of mutual node fr…
In (exploratory) factor analysis, the loading matrix is identified only up to orthogonal rotation. For identifiability, one thus often takes the loading matrix to be lower triangular with positive diagonal entries. In Bayesian inference, a standard practice is then to specify a prior under which the loadings are indepe…
Paper solves a key problem in learning from high-dimensional covariance matrices.
Undirected graphs can be used to describe matrix variate distributions. In this paper, we develop new methods for estimating the graphical structures and underlying parameters, namely, the row and column covariance and inverse covariance matrices from the matrix variate data. Under sparsity conditions, we show that one…
Paper improves matrix-valued data classification using nonparametric LDA.
CoreFlow models matrix-valued distributions efficiently, preserving shared low-rank structure.
In this note we answer a question of G. Lecué, by showing that column normalization of a random matrix with iid entries need not lead to good sparse recovery properties, even if the generating random variable has a reasonable moment growth. Specifically, for every we construct a random vector …
Muon dynamics study uses spectral Wasserstein flow for optimization stability.
The paper addresses statistical inference in matching markets with dependent missingness.
The paper improves matrix completion with auxiliary covariates using LS estimation.
New method corrects bias in missing data for matrix completion.
Paper proposes an online estimator for covariance matrix of SGD iterates.
Nonnegative matrix factorization (NMF) has been shown recently to be tractable under the separability assumption, under which all the columns of the input data matrix belong to the convex cone generated by only a few of these columns. Bittorf, Recht, Ré and Tropp (`Factoring nonnegative matrices with linear programs', …
Missing data estimation is an important challenge with high-dimensional data arranged in the form of a matrix. Typically this data matrix is transposable, meaning that either the rows, columns or both can be treated as features. To model transposable data, we present a modification of the matrix-variate normal, the mea…
This paper is on the normal approximation of singular subspaces when the noise matrix has i.i.d. entries. Our contributions are three-fold. First, we derive an explicit representation formula of the empirical spectral projectors. The formula is neat and holds for deterministic matrix perturbations. Second, we calculate…
In this paper, we use a new approach to prove that the largest eigenvalue of the sample covariance matrix of a normally distributed vector is bigger than the true largest eigenvalue with probability 1 when the dimension is infinite. We prove a similar result for the smallest eigenvalue.
Several variants of recurrent neural networks (RNNs) with orthogonal or unitary recurrent matrices have recently been developed to mitigate the vanishing/exploding gradient problem and to model long-term dependencies of sequences. However, with the eigenvalues of the recurrent matrix on the unit circle, the recurrent s…
The paper develops inference methods for high-dimensional multi-task regression with row-sparse coefficients.
New optimizers control network width scaling, improving stability and transfer across different model sizes.
We are concerned with an approximation problem for a symmetric positive semidefinite matrix due to motivation from a class of nonlinear machine learning methods. We discuss an approximation approach that we call {matrix ridge approximation}. In particular, we define the matrix ridge approximation as an incomplete matri…
This paper is a further extension of the method proposed in Itkin, 2014 as applied to another set of jump-diffusion models: Inverse Normal Gaussian, Hyperbolic and Meixner. To solve the corresponding PIDEs we accomplish few steps. First, a second-order operator splitting on financial processes (diffusion and jumps) is …
Improved bipartite link prediction using 2-hop paths.
Low-rank matrix regression refers to the instances of recovering a low-rank matrix based on specially designed measurements and the corresponding noisy outcomes. In the last decade, numerous statistical methodologies have been developed for efficiently recovering the unknown low-rank matrices. However, in some applicat…
We investigate the analogy between the large N expansion in normal matrix models and the asymptotic expansion of the determinant of the Hilb map, appearing in the study of critical metrics on complex manifolds via projective embeddings. This analogy helps to understand the geometric meaning of the expansion of matrix m…
New connections on symmetric spaces with invariant properties.
Study on neural networks with non-normal interactions reveals unique spectral properties.
MANO normalizes logits to estimate test accuracy without labels.
The Sinkhorn-Knopp algorithm converges quickly but the number of iterations is poorly understood.
A new method for learning Bayesian neural networks using layerwise inference.