Simple bounds for covariance and Gram matrices across various settings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep kernel processes unify various models using Gram matrices and kernel functions.
When presented with Out-of-Distribution (OOD) examples, deep neural networks yield confident, incorrect predictions. Detecting OOD examples is challenging, and the potential risks are high. In this paper, we propose to detect OOD examples by identifying inconsistencies between activity patterns and class predicted. We …
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
We compute Stokes matrices and monodromy for the quantum cohomology of projective spaces. We prove that the Stokes' matrix of the quantum cohomology coincides with the Gram matrix in the theory of derived categories of coherent sheaves.
Given a real matrix A with n columns, the problem is to approximate the Gram product AA^T by c << n weighted outer products of columns of A. Necessary and sufficient conditions for the exact computation of AA^T (in exact arithmetic) from c >= rank(A) columns depend on the right singular vector matrix of A. For a Monte-…
New model for multi-layer categorical data improves latent class analysis.
We establish the second part of Milnor's conjecture on the volume of simplexes in hyperbolic and spherical spaces. A characterization of the closure of the space of the angle Gram matrices of simplexes is also obtained.
We present power low rank ensembles (PLRE), a flexible framework for n-gram language modeling where ensembles of low rank matrices and tensors are used to obtain smoothed probability estimates of words in context. Our method can be understood as a generalization of n-gram modeling to non-integer n, and includes standar…
This paper provides a generic framework of component analysis (CA) methods introducing a new expression for scatter matrices and Gram matrices, called Generalized Pairwise Expression (GPE). This expression is quite compact but highly powerful: The framework includes not only (1) the standard CA methods but also (2) sev…
Many interesting machine learning problems are best posed by considering instances that are distributions, or sample sets drawn from distributions. Previous work devoted to machine learning tasks with distributional inputs has done so through pairwise kernel evaluations between pdfs (or sample sets). While such an appr…
New method improves feasibility of fitting Gaussian vectors to an ellipsoid.
Estimates latent norms and Gram matrices for graphs on Euclidean balls.
Neuc-MDS extends MDS for non-Euclidean data.
Kernel matrices (e.g. Gram or similarity matrices) are essential for many state-of-the-art approaches to classification, clustering, and dimensionality reduction. For large datasets, the cost of forming and factoring such kernel matrices becomes intractable. To address this challenge, we introduce a new adaptive sampli…
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural network…
Improved variational approximation for deep Wishart process models.
Spectral graph sparsification preserves geometry of GNN embeddings.
The Gram determinant of type was introduced by Lickorish in his work on invariants of 3 - manifolds. We generalize the theory of the Gram determinant of type by evaluating, in the annulus, a bilinear form of non-intersecting connections in the disc. The main result provides a closed formula for this Gram determ…
We study the energy distribution of harmonic 1-forms on a compact hyperbolic Riemann surface where a short closed geodesic is pinched. If the geodesic separates the surface into two parts, then the Jacobian torus of develops into a torus that splits. If the geodesic is nonseparating then the Jacobian torus of $…
Eliciting semantic similarity between concepts in the biomedical domain remains a challenging task. Recent approaches founded on embedding vectors have gained in popularity as they risen to efficiently capture semantic relationships The underlying idea is that two words that have close meaning gather similar contexts. …
We investigate the Gram determinant of the bilinear form based on curves in a planar surface, with a focus on the disk with two holes. We prove that the determinant based on curves divides the determinant based on curves. Motivated by the work on Gram determinants based on curves in a disk and curves in an an…
Study Gram determinants in knot theory, focusing on a Möbius band determinant.
The paper connects Chebyshev polynomials and Gram determinants on Möbius bands.
Proposes a method to infer complex network topologies from multiple graphs.
Any-gram kernels are a flexible and efficient way to employ bag-of-n-gram features when learning from textual data. They are also compatible with the use of word embeddings so that word similarities can be accounted for. While the original any-gram kernels are implemented on top of tree kernels, we propose a new approa…
A new method for deep Wishart processes improves kernel-based models.
Effective Gram matrix predicts deep network generalization.
Method reduces categorical data to lower dimensions using density matrices.
This work presents a parametrized family of divergences, namely Alpha-Beta Log- Determinant (Log-Det) divergences, between positive definite unitized trace class operators on a Hilbert space. This is a generalization of the Alpha-Beta Log-Determinant divergences between symmetric, positive definite matrices to the infi…
New Gram determinant from Möbius band connects to annulus case.
Existing multi-view learning methods based on kernel function either require the user to select and tune a single predefined kernel or have to compute and store many Gram matrices to perform multiple kernel learning. Apart from the huge consumption of manpower, computation and memory resources, most of these models see…
Deep learning methods exhibit promising performance for predictive modeling in healthcare, but two important challenges remain: -Data insufficiency:Often in healthcare predictive modeling, the sample size is insufficient for deep learning methods to achieve satisfactory results. -Interpretation:The representations lear…
We present network embedding algorithms that capture information about a node from the local distribution over node attributes around it, as observed over random walks following an approach similar to Skip-gram. Observations from neighborhoods of different sizes are either pooled (AE) or encoded distinctly in a multi-s…
Corrected CBOW performs similarly to Skip-gram.
Researchers use quantum chaos and RMT to analyze turbulence, revealing unique scaling laws.
New method corrects missing data bias in dimension reduction.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
N-grams have been a common tool for information retrieval and machine learning applications for decades. In nearly all previous works, only a few values of are tested, with being exceedingly rare. Larger values of are not tested due to computational burden or the fear of overfitting. In this work, we pr…
Motivated by the fact that most of the information relevant to the prediction of target tokens is drawn from the source sentence , we propose truncating the target-side window used for computing self-attention by making an -gram assumption. Experiments on WMT EnDe and EnFr data sets show that the…
We use the Jones-Wenzl idempotents to construct a basis of Temperley-Lieb algebra TL_n. This allows a short calculation for a Gram determinant of Lickorish's bilinear form on the Temperley-Lieb algebra.
We present in this work a new family of kernels to compare positive measures on arbitrary spaces $\Xcal$ endowed with a positive kernel , which translates naturally into kernels between histograms or clouds of points. We first cover the case where $\Xcal$ is Euclidian, and focus on kernels which take into account th…
New metric learning approach for tree data reduces computation cost.
In the modern age, rankings data is ubiquitous and it is useful for a variety of applications such as recommender systems, multi-object tracking and preference learning. However, most rankings data encountered in the real world is incomplete, which prevents the direct application of existing modelling tools for complet…
Gradient descent aligns neural feature matrices with pre-activation tangent features.
We show that the skip-gram formulation of word2vec trained with negative sampling is equivalent to a weighted logistic PCA. This connection allows us to better understand the objective, compare it to other word embedding methods, and extend it to higher dimensional models.
Here we present a novel approach to statistical analysis of financial time series. The approach is based on -grams frequency dictionaries derived from the quantized market data. Such dictionaries are studied by evaluating their information capacity using relative entropy. A specific quantization of (originally conti…
Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.