The Procrustes distance is used to quantify the similarity or dissimilarity of (3-dimensional) shapes, and extensively used in biological morphometrics. Typically each (normalized) shape is represented by N landmark points, chosen to be homologous (i.e. corresponding to each other), as far as possible, and the Procrust…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method for computing shape barycenters from point clouds using Procrustes-Wasserstein distance.
Extends metrics for SPD matrices to infinite dimensions.
New measures link neural representation geometry to decoding ability.
Modified Wasserstein metric for Gaussian distributions, invariant to isometries.
The paper introduces metrics for robust unsupervised learning of vehicle interactions.
LDLE embeds manifolds in lower dimensions with low distortion.
We solve a key problem in cross-lingual learning using a novel approach.
Method preserves correlations in synthetic data.
Python package for SPD matrix distances, reproducible and extensible.
We present the Procrustes measure, a novel measure based on Procrustes rotation that enables quantitative comparison of the output of manifold-based embedding algorithms (such as LLE (Roweis and Saul, 2000) and Isomap (Tenenbaum et al, 2000)). The measure also serves as a natural tool when choosing dimension-reduction …
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
Multivariate Analysis (MVA) comprises a family of well-known methods for feature extraction that exploit correlations among input variables of the data representation. One important property that is enjoyed by most such methods is uncorrelation among the extracted features. Recently, regularized versions of MVA methods…
Paper proposes a KGE framework that reduces training time and carbon footprint.
Representational similarity analysis (RSA) has been shown to be an effective framework to characterize brain-activity profiles and deep neural network activations as representational geometry by computing the pairwise distances of the response patterns as a representational dissimilarity matrix (RDM). However, how to p…
One of the common tasks in unsupervised learning is dimensionality reduction, where the goal is to find meaningful low-dimensional structures hidden in high-dimensional data. Sometimes referred to as manifold learning, this problem is closely related to the problem of localization, which aims at embedding a weighted gr…
Study matches two noisy point clouds with geometric transformations and relabeling.
Solves a fundamental problem in statistics and imaging with new methods.
Algorithm recovers permutations of high-dimensional Gaussian vectors with constant correlation.
A new algorithm reduces communication in distributed SVD by factors.
A new method for aligning datasets without known correspondences.
Mapping and translating professional but arcane clinical jargons to consumer language is essential to improve the patient-clinician communication. Researchers have used the existing biomedical ontologies and consumer health vocabulary dictionary to translate between the languages. However, such approaches are limited b…
Deep model predicts shapes of curves with multiple covariates.
The problem of estimating sparse eigenvectors of a symmetric matrix attracts a lot of attention in many applications, especially those with high dimensional data set. While classical eigenvectors can be obtained as the solution of a maximization problem, existing approaches formulate this problem by adding a penalty te…
We have observed an interesting, yet unexplained, phenomenon: Semidefinite programming (SDP) based relaxations of maximum likelihood estimators (MLE) tend to be tight in recovery problems with noisy data, even when MLE cannot exactly recover the ground truth. Several results establish tightness of SDP based relaxations…
Gradient descent solves asymmetric low-rank matrix sensing without balancing.
Geometric stability predicts steerability and detects drift in language models.
Proposes BONMI for integrating noisy matrices from multi-source data.
Change detection in dynamic networks is an important problem in many areas, such as fraud detection, cyber intrusion detection and health care monitoring. It is a challenging problem because it involves a time sequence of graphs, each of which is usually very large and sparse with heterogeneous vertex degrees, resultin…
Survey of Locally Linear Embedding and its variants.
Robustly computes intrinsic coordinates on point clouds using resampling and averaging.
Geometric stability measures neural network robustness, distinguishing from similarity metrics.
We revisit the inductive matrix completion problem that aims to recover a rank- matrix with ambient dimension given features as the side prior information. The goal is to make use of the known features to reduce sample and computational complexities. We present and analyze a new gradient-based non-convex…
Proposes a method to align language and image data.
Study finds whitepaper narratives do not predict market factor structure.
We consider the task of aligning two sets of points in high dimension, which has many applications in natural language processing and computer vision. As an example, it was recently shown that it is possible to infer a bilingual lexicon, without supervised data, by aligning word embeddings trained on monolingual data. …
We propose a data aggregation-based algorithm with monotonic convergence to a global optimum for a generalized version of the L1-norm error fitting model with an assumption of the fitting function. The proposed algorithm generalizes the recent algorithm in the literature, aggregate and iterative disaggregate (AID), whi…
FedPower improves eigenspace estimation privacy in federated learning.
New metrics defined on SPD matrices link to divergences and curvature.
New algorithm for robust circular coordinates in recurrent time series data.
New method preserves privacy while detecting communities in distributed networks.
The paper tightens bounds on distances between Reeb graphs.
Extends Teichmüller distance concept to non-distance maps.
The Wasserstein distance and its variations, e.g., the sliced-Wasserstein (SW) distance, have recently drawn attention from the machine learning community. The SW distance, specifically, was shown to have similar properties to the Wasserstein distance, while being much simpler to compute, and is therefore used in vario…
We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new distances are metrics whenever the ground distance between the marginals is a metric, generalize both…
In the present paper we calculate the Gromov-Hausdorff distance between an arbitrary simplex (a metric space all whose non-zero distances are the same) and a finite metric space whose non-zero distances take two distinct values (so-called -distance spaces). As a corollary, a complete solution to generalized Borsuk p…
New toolkit for directed distances improves flexibility of OT problems.
There have lately been several suggestions for parametrized distances on a graph that generalize the shortest path distance and the commute time or resistance distance. The need for developing such distances has risen from the observation that the above-mentioned common distances in many situations fail to take into ac…