A method for learning embeddings from multi-view data using Gromov-Wasserstein.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Two algorithms estimate Wasserstein distance matrices from few entries for manifold learning.
Optimizes embedding accuracy for data variance and error.
Extends graph encoder embedding to weighted graphs and matrices.
Two methods factor out prior knowledge from low-dimensional embeddings.
New theory for eigenvectors of generalized Laplacian matrices, addressing dependency issues.
Embedding complex objects as vectors in low dimensional spaces is a longstanding problem in machine learning. We propose in this work an extension of that approach, which consists in embedding objects as elliptical probability distributions, namely distributions whose densities have elliptical level sets. We endow thes…
We present a new paradigm for speeding up randomized computations of several frequently used functions in machine learning. In particular, our paradigm can be applied for improving computations of kernels based on random embeddings. Above that, the presented framework covers multivariate randomized functions. As a bypr…
Paper introduces a new distance measure for Gaussian Mixture Models.
Flexible embedding framework for diverse data types.
The fields of compressed sensing (CS) and matrix completion have shown that high-dimensional signals with sparse or low-rank structure can be effectively projected into a low-dimensional space (for efficient acquisition or processing) when the projection operator achieves a stable embedding of the data by satisfying th…
We examine a class of embeddings based on structured random matrices with orthogonal rows which can be applied in many machine learning applications including dimensionality reduction and kernel approximation. For both the Johnson-Lindenstrauss transform and the angular kernel, we show that we can select matrices yield…
EPINE enhances network embedding by improving adjacency matrix-based high-order proximity.
In statistical relational learning, knowledge graph completion deals with automatically understanding the structure of large knowledge graphs---labeled directed graphs---and predicting missing relationships---labeled edges. State-of-the-art embedding models propose different trade-offs between modeling expressiveness, …
CMF is a technique for simultaneously learning low-rank representations based on a collection of matrices with shared entities. A typical example is the joint modeling of user-item, item-property, and user-feature matrices in a recommender system. The key idea in CMF is that the embeddings are shared across the matrice…
Optimal subspace embedding with near-optimal sparsity for high-dimensional data.
Recent advances suggest that encoding images through Symmetric Positive Definite (SPD) matrices and then interpreting such matrices as points on Riemannian manifolds can lead to increased classification performance. Taking into account manifold geometry is typically done via (1) embedding the manifolds in tangent space…
Study of discrete period matrices on embedded graphs, relating to Riemann surfaces.
Linear representations help embed manifolds into matrix spaces.
Develops log-Euclidean Lie groups for SPD and correlation matrices.
Isometry pursuit identifies orthonormal submatrices from wide matrices.
Researchers develop geodesics for a new metric on correlation matrices.
Geometric approach for unsupervised word embedding alignment.
Feature extraction and dimension reduction for networks is critical in a wide variety of domains. Efficiently and accurately learning features for multiple graphs has important applications in statistical inference on graphs. We propose a method to jointly embed multiple undirected graphs. Given a set of graphs, the jo…
In this paper we show that for the purposes of dimensionality reduction certain class of structured random matrices behave similarly to random Gaussian matrices. This class includes several matrices for which matrix-vector multiply can be computed in log-linear time, providing efficient dimensionality reduction of gene…
Unified framework for hyperbolic embeddings from mixed data types.
We present a theory for Euclidean dimensionality reduction with subgaussian matrices which unifies several restricted isometry property and Johnson-Lindenstrauss type results obtained earlier for specific data sets. In particular, we recover and, in several cases, improve results for sets of sparse and structured spars…
The abstract theorem is extended to higher genus surfaces.
The space of matrices of positive determinant GL^+_n inherits an extrinsic metric space structure from R^{n^2}. On the other hand, taking the infimum of the lengths of all paths connecting two points in GL^+_n gives an intrinsic metric. We prove bilipschitz equivalence for intrinsic and extrinsic metrics on GL^+_n, exp…
Model compression is essential for serving large deep neural nets on devices with limited resources or applications that require real-time responses. As a case study, a state-of-the-art neural language model usually consists of one or more recurrent layers sandwiched between an embedding layer used for representing inp…
Method reduces categorical data to lower dimensions using density matrices.
We propose a scheme for recycling Gaussian random vectors into structured matrices to approximate various kernel functions in sublinear time via random embeddings. Our framework includes the Fastfood construction as a special case, but also extends to Circulant, Toeplitz and Hankel matrices, and the broader family of s…
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
The paper presents two schemes for sampling matrices from specific distributions on a manifold.
SpecNet2 improves spectral embedding without orthogonalization, achieving better performance and efficiency.
Proposes a new graph representation method using tensor products.
We prove, using the subspace embedding guarantee in a black box way, that one can achieve the spectral norm guarantee for approximate matrix multiplication with a dimensionality-reducing map having rows. Here is the maximum stable rank, i.e. squared ratio of Frobenius and op…
Generating point clouds, e.g., molecular structures, in arbitrary rotations, translations, and enumerations remains a challenging task. Meanwhile, neural networks utilizing symmetry invariant layers have been shown to be able to optimize their training objective in a data-efficient way. In this spirit, we present an ar…
In this paper, we study the problem of approximately computing the product of two real matrices. In particular, we analyze a dimensionality-reduction-based approximation algorithm due to Sarlos [1], introducing the notion of nuclear rank as the ratio of the nuclear norm over the spectral norm. The presented bound has i…
Investigates O(n)-invariant metrics on SPD matrices, extending kernel metrics.
Improved bounds for sensitivity sampling reducing the sample complexity for structured matrices.
This paper deals with two related problems, namely distance-preserving binary embeddings and quantization for compressed sensing . First, we propose fast methods to replace points from a subset , associated with the Euclidean metric, with points in the cube and we associa…
Sparse oblique decision tree improves security rules for renewable power systems.
Estimates low-rank distributional matrices from incomplete samples.
This paper considers the problem of embedding directed graphs in Euclidean space while retaining directional information. We model a directed graph as a finite set of observations from a diffusion on a manifold endowed with a vector field. This is the first generative model of its kind for directed graphs. We introduce…
Proteins are the major building blocks of life, and actuators of almost all chemical and biophysical events in living organisms. Their native structures in turn enable their biological functions which have a fundamental role in drug design. This motivates predicting the structure of a protein from its sequence of amino…
This paper considers *-graphs in which all vertices have degree 4 or 6, and studies the question of calculating the genus of nonorientable surfaces into which such graphs may be embedded. In a previous paper by the authors, the problem of calculating whether a given *-graph in which all vertices have degree 4 or 6 admi…
Proposes a new algorithm to estimate invariant subspaces across multilayer networks.