A clustering algorithm uses the left Gram matrix for high dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Effective Gram matrix predicts deep network generalization.
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
New method corrects missing data bias in dimension reduction.
Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.
Deep kernel processes unify various models using Gram matrices and kernel functions.
This paper extends the convergence rate of DEQs with ReLU to any general activation.
The paper connects Chebyshev polynomials and Gram determinants on Möbius bands.
Paper proposes a method to recover point configurations from noisy distance data.
To any compact Riemann surface of genus g one may assign a principally polarized abelian variety of dimension g, the Jacobian of the Riemann surface. The Jacobian is a complex torus, and a Gram matrix of the lattice of a Jacobian is called a period Gram matrix. This paper provides upper and lower bounds for all the ent…
Given a real matrix A with n columns, the problem is to approximate the Gram product AA^T by c << n weighted outer products of columns of A. Necessary and sufficient conditions for the exact computation of AA^T (in exact arithmetic) from c >= rank(A) columns depend on the right singular vector matrix of A. For a Monte-…
We compute Stokes matrices and monodromy for the quantum cohomology of projective spaces. We prove that the Stokes' matrix of the quantum cohomology coincides with the Gram matrix in the theory of derived categories of coherent sheaves.
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero training loss in polynomial time for a deep over-parameterized neural network with residual connections (ResNet). Our analysis relies on the p…
Since the invention of word2vec, the skip-gram model has significantly advanced the research of network embedding, such as the recent emergence of the DeepWalk, LINE, PTE, and node2vec approaches. In this work, we show that all of the aforementioned models with negative sampling can be unified into the matrix factoriza…
Estimates latent norms and Gram matrices for graphs on Euclidean balls.
This paper deals with the design of a sensing matrix along with a sparse recovery algorithm by utilizing the probability-based prior information for compressed sensing system. With the knowledge of the probability for each atom of the dictionary being used, a diagonal weighted matrix is obtained and then the sensing ma…
New estimator stabilizes higher-order influence functions for stable statistical inference.
Mixed-precision CA-SGD for generalized linear models on GPUs
A new method uses Gram matrix for efficient multivariate functional principal components.
Many interesting machine learning problems are best posed by considering instances that are distributions, or sample sets drawn from distributions. Previous work devoted to machine learning tasks with distributional inputs has done so through pairwise kernel evaluations between pdfs (or sample sets). While such an appr…
When presented with Out-of-Distribution (OOD) examples, deep neural networks yield confident, incorrect predictions. Detecting OOD examples is challenging, and the potential risks are high. In this paper, we propose to detect OOD examples by identifying inconsistencies between activity patterns and class predicted. We …
The paper studies geometric structures on SL(n,R) induced by the Killing form.
A new matrix concentration inequality for random products of matrices.
We present in this work a new family of kernels to compare positive measures on arbitrary spaces $\Xcal$ endowed with a positive kernel , which translates naturally into kernels between histograms or clouds of points. We first cover the case where $\Xcal$ is Euclidian, and focus on kernels which take into account th…
A new method learns dynamic graph representations from time-varying data.
This paper analyzes error in SKI for Gaussian Processes, providing conditions for linear time inference.
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural network…
Kernel methods are ubiquitous tools in machine learning. However, there is often little reason for the common practice of selecting a kernel a priori. Even if a universal approximating kernel is selected, the quality of the finite sample estimator may be greatly affected by the choice of kernel. Furthermore, when direc…
It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings. Recent papers have proved this statement theoretically for highly overparametrized networks under reasonable assumptions. These results either assume…
We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
Researchers use quantum chaos and RMT to analyze turbulence, revealing unique scaling laws.
Asymptotics of quantum symbols corresponding to a hyperbolic tetrahedra is investigated and the first two leading terms are determined for the case that the tetrahedron has a ideal or ultra-ideal vertex. These terms are given by the volume and the determinant of the Gram matrix of the tetrahedron. A relation to th…
We uncover scaling laws and statistical structure in complex datasets.
The Gram determinant of type was introduced by Lickorish in his work on invariants of 3 - manifolds. We generalize the theory of the Gram determinant of type by evaluating, in the annulus, a bilinear form of non-intersecting connections in the disc. The main result provides a closed formula for this Gram determ…
New estimator stabilizes higher-order influence functions for bilinear forms.
Eliciting semantic similarity between concepts in the biomedical domain remains a challenging task. Recent approaches founded on embedding vectors have gained in popularity as they risen to efficiently capture semantic relationships The underlying idea is that two words that have close meaning gather similar contexts. …
Random feature maps are ubiquitous in modern statistical machine learning, where they generalize random projections by means of powerful, yet often difficult to analyze nonlinear operators. In this paper, we leverage the "concentration" phenomenon induced by random matrix theory to perform a spectral analysis on the Gr…
We investigate the Gram determinant of the bilinear form based on curves in a planar surface, with a focus on the disk with two holes. We prove that the determinant based on curves divides the determinant based on curves. Motivated by the work on Gram determinants based on curves in a disk and curves in an an…
Neuc-MDS extends MDS for non-Euclidean data.
Study Gram determinants in knot theory, focusing on a Möbius band determinant.
Any-gram kernels are a flexible and efficient way to employ bag-of-n-gram features when learning from textual data. They are also compatible with the use of word embeddings so that word similarities can be accounted for. While the original any-gram kernels are implemented on top of tree kernels, we propose a new approa…
Linearized attention fails to converge to NTK limit even at large widths.
Study spectral properties of sparse random graphs to recover latent vectors.
New Gram determinant from Möbius band connects to annulus case.
Spectral graph sparsification preserves geometry of GNN embeddings.
The Information Plane theory predicts autoencoders do not compress input information.
Method reduces categorical data to lower dimensions using density matrices.