A clustering algorithm uses the left Gram matrix for high dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Effective Gram matrix predicts deep network generalization.
New estimator stabilizes higher-order influence functions for stable statistical inference.
Estimates latent norms and Gram matrices for graphs on Euclidean balls.
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
New method corrects missing data bias in dimension reduction.
To any compact Riemann surface of genus g one may assign a principally polarized abelian variety of dimension g, the Jacobian of the Riemann surface. The Jacobian is a complex torus, and a Gram matrix of the lattice of a Jacobian is called a period Gram matrix. This paper provides upper and lower bounds for all the ent…
A new method uses Gram matrix for efficient multivariate functional principal components.
Many interesting machine learning problems are best posed by considering instances that are distributions, or sample sets drawn from distributions. Previous work devoted to machine learning tasks with distributional inputs has done so through pairwise kernel evaluations between pdfs (or sample sets). While such an appr…
New estimator stabilizes higher-order influence functions for bilinear forms.
Deep kernel processes unify various models using Gram matrices and kernel functions.
This paper extends the convergence rate of DEQs with ReLU to any general activation.
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural network…
The paper connects Chebyshev polynomials and Gram determinants on Möbius bands.
Paper proposes a method to recover point configurations from noisy distance data.
Given a real matrix A with n columns, the problem is to approximate the Gram product AA^T by c << n weighted outer products of columns of A. Necessary and sufficient conditions for the exact computation of AA^T (in exact arithmetic) from c >= rank(A) columns depend on the right singular vector matrix of A. For a Monte-…
This paper analyzes error in SKI for Gaussian Processes, providing conditions for linear time inference.
We compute Stokes matrices and monodromy for the quantum cohomology of projective spaces. We prove that the Stokes' matrix of the quantum cohomology coincides with the Gram matrix in the theory of derived categories of coherent sheaves.
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero training loss in polynomial time for a deep over-parameterized neural network with residual connections (ResNet). Our analysis relies on the p…
Since the invention of word2vec, the skip-gram model has significantly advanced the research of network embedding, such as the recent emergence of the DeepWalk, LINE, PTE, and node2vec approaches. In this work, we show that all of the aforementioned models with negative sampling can be unified into the matrix factoriza…
Kernel methods are ubiquitous tools in machine learning. However, there is often little reason for the common practice of selecting a kernel a priori. Even if a universal approximating kernel is selected, the quality of the finite sample estimator may be greatly affected by the choice of kernel. Furthermore, when direc…
The Information Plane theory predicts autoencoders do not compress input information.
Paper proposes detecting OOD examples using Gram matrices and in-distribution data.
Method reduces categorical data to lower dimensions using density matrices.
This paper deals with the design of a sensing matrix along with a sparse recovery algorithm by utilizing the probability-based prior information for compressed sensing system. With the knowledge of the probability for each atom of the dictionary being used, a diagonal weighted matrix is obtained and then the sensing ma…
Mixed-precision CA-SGD for generalized linear models on GPUs
New method stabilizes private LASSO for high-dimensional data with diverse covariate scales.
Proposes a method to infer complex network topologies from multiple graphs.
Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…
We present in this work a new family of kernels to compare positive measures on arbitrary spaces $\Xcal$ endowed with a positive kernel , which translates naturally into kernels between histograms or clouds of points. We first cover the case where $\Xcal$ is Euclidian, and focus on kernels which take into account th…
A new method learns dynamic graph representations from time-varying data.
This paper is concerned with estimating the column space of an unknown low-rank matrix , given noisy and partial observations of its entries. There is no shortage of scenarios where the observations -- while being too noisy to support faithful recovery of the ent…
It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings. Recent papers have proved this statement theoretically for highly overparametrized networks under reasonable assumptions. These results either assume…
Characterizes RFF regression in large setting, providing precise learning phases and double descent curve.
Most machine learning algorithms, such as classification or regression, treat the individual data point as the object of interest. Here we consider extending machine learning algorithms to operate on groups of data points. We suggest treating a group of data points as an i.i.d. sample set from an underlying feature dis…
We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
Existing multi-view learning methods based on kernel function either require the user to select and tune a single predefined kernel or have to compute and store many Gram matrices to perform multiple kernel learning. Apart from the huge consumption of manpower, computation and memory resources, most of these models see…
Researchers use quantum chaos and RMT to analyze turbulence, revealing unique scaling laws.
Asymptotics of quantum symbols corresponding to a hyperbolic tetrahedra is investigated and the first two leading terms are determined for the case that the tetrahedron has a ideal or ultra-ideal vertex. These terms are given by the volume and the determinant of the Gram matrix of the tetrahedron. A relation to th…
We uncover scaling laws and statistical structure in complex datasets.
The Gram determinant of type was introduced by Lickorish in his work on invariants of 3 - manifolds. We generalize the theory of the Gram determinant of type by evaluating, in the annulus, a bilinear form of non-intersecting connections in the disc. The main result provides a closed formula for this Gram determ…
Word embeddings learnt from large corpora have been adopted in various applications in natural language processing and served as the general input representations to learning systems. Recently, a series of post-processing methods have been proposed to boost the performance of word embeddings on similarity comparison an…
Proposes SNML for selecting word2vec Skip-gram dimensionality.
Random feature maps are ubiquitous in modern statistical machine learning, where they generalize random projections by means of powerful, yet often difficult to analyze nonlinear operators. In this paper, we leverage the "concentration" phenomenon induced by random matrix theory to perform a spectral analysis on the Gr…
Eliciting semantic similarity between concepts in the biomedical domain remains a challenging task. Recent approaches founded on embedding vectors have gained in popularity as they risen to efficiently capture semantic relationships The underlying idea is that two words that have close meaning gather similar contexts. …
We investigate the Gram determinant of the bilinear form based on curves in a planar surface, with a focus on the disk with two holes. We prove that the determinant based on curves divides the determinant based on curves. Motivated by the work on Gram determinants based on curves in a disk and curves in an an…
Neuc-MDS extends MDS for non-Euclidean data.