Manifolds uniquely identified by boundary distance differences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes how the one-dimensional Wasserstein distance captures pointwise density differences in finite samples.
The study compares Euclidean and cosine distances in medical drug prescription prediction.
A framework for measuring differences in categorical data.
CADM proposes a cluster-specific distance metric for categorical data clustering.
We propose unsupervised representation learning and feature extraction from dendrograms. The commonly used Minimax distance measures correspond to building a dendrogram with single linkage criterion, with defining specific forms of a level function and a distance function over that. Therefore, we extend this method to …
As a fundamental problem of natural language processing, it is important to measure the distance between different documents. Among the existing methods, the Word Mover's Distance (WMD) has shown remarkable success in document semantic matching for its clear physical insight as a parameter-free model. However, WMD is e…
This work examines the sensitivity of energy distance to mean differences compared to covariance differences.
Causal inference relies on the structure of a graph, often a directed acyclic graph (DAG). Different graphs may result in different causal inference statements and different intervention distributions. To quantify such differences, we propose a (pre-) distance between DAGs, the structural intervention distance (SID). T…
This guide explains statistical distances for evaluating generative models.
A new metric compares true and learned causal graphs considering data and graph structure.
Learning algorithms for implicit generative models can optimize a variety of criteria that measure how the data distribution differs from the implicit model distribution, including the Wasserstein distance, the Energy distance, and the Maximum Mean Discrepancy criterion. A careful look at the geometries induced by thes…
The distribution is able to characterize different regions in monopolarized SAR imagery. It is indexed by three parameters: the number of looks (which can be estimated in the whole image), a scale parameter and a texture parameter. This paper presents a new proposal for feature extraction and region d…
Physics: Similar long-distance properties can mask vastly different short-distance metrics.
In high dimension, low sample size (HDLSS)settings, the simple average distance classifier based on the Euclidean distance performs poorly if differences between the locations get masked by the scale differences. To rectify this issue, modifications to the average distance classifier was proposed by Chan and Hall (2009…
An alternating distance is a link invariant that measures how far away a link is from alternating. We study several alternating distances and demonstrate that there exist families of links for which the difference between certain alternating distances is arbitrarily large. We also show that two alternating distances, t…
New measures assess differences in causal graphs' separations.
We show that on a certain class of bounded, complete Reinhardt domains in that enjoy a lot of symmetries, the Carathéodory pseudo-distance and the geodesic distance of the complete Kähler-Einstein metric with Ricci curvature are different.
New examples show flip distance and polyhedron triangulation numbers differ, with ratio close to 3/2.
New distances for causal graphs improve evaluation of learned structures.
In this paper we tackle the issue of clustering trajectories of geolocalized observations. Using clustering technics based on the choice of a distance between the observations, we first provide a comprehensive review of the different distances used in the literature to compare trajectories. Then based on the limitation…
A new robust metric compares distributions more accurately than existing methods.
Generative Adversarial Networks (GANs) have been used to model the underlying probability distribution of sample based datasets. GANs are notoriuos for training difficulties and their dependence on arbitrary hyperparameters. One recent improvement in GAN literature is to use the Wasserstein distance as loss function le…
Diffusion-weighted MR imaging (DWI) is the only method we currently have to measure connections between different parts of the human brain in vivo. To elucidate the structure of these connections, algorithms for tracking bundles of axonal fibers through the subcortical white matter rely on local estimates of the fiber …
We develop a novel methodology based on the marriage between the Bhattacharyya distance, a measure of similarity across distributions of random variables, and the Johnson-Lindenstrauss Lemma, a technique for dimension reduction. The resulting technique is a simple yet powerful tool that allows comparisons between data-…
Paper uses Gaussian mixture models and Wasserstein distance for schema matching.
Tractograms are mathematical representations of the main paths of axons within the white matter of the brain, from diffusion MRI data. Such representations are in the form of polylines, called streamlines, and one streamline approximates the common path of tens of thousands of axons. The analysis of tractograms is a ta…
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For …
Metric learning has the aim to improve classification accuracy by learning a distance measure which brings data points from the same class closer together and pushes data points from different classes further apart. Recent research has demonstrated that metric learning approaches can also be applied to trees, such as m…
Region-based classification of PolSAR data can be effectively performed by seeking for the assignment that minimizes a distance between prototypes and segments. Silva et al (2013) used stochastic distances between complex multivariate Wishart models which, differently from other measures, are computationally tractable.…
Proposes HOT method for robust multi-view learning.
Efficient sampling reduces memory usage for Minimax distance analysis.
Metric learning has the aim to improve classification accuracy by learning a distance measure which brings data points from the same class closer together and pushes data points from different classes further apart. Recent research has demonstrated that metric learning approaches can also be applied to trees, such as m…
The paper uses distance correlation for brain connectivity and a novel multi-task learning model for age prediction.
Wasserstein GANs fail to approximate Wasserstein distance, leading to their success.
This study improves graph coarsening methods by preserving graph spectrum and distances.
We introduce a new distance metric for non-linear embeddings of Tempered Exponential Measures.
Graph distance metric learning serves as the foundation for many graph learning problems, e.g., graph clustering, graph classification and graph matching. Existing research works on graph distance metric (or graph kernels) learning fail to maintain the basic properties of such metrics, e.g., non-negative, identity of i…
New method uses spherical convolutional Wasserstein distance to validate climate models.
A new method compares unaligned datasets using log-Euclidean signatures of SPD matrices.
The article proposes modified Gower's coefficients for handling mixed type variables in nearest neighbor methods.
Exploring distance functions on spacetime models.
For localization and mapping of indoor environments through WiFi signals, locations are often represented as likelihoods of the received signal strength indicator. In this work we compare various measures of distance between such likelihoods in combination with different methods for estimation and representation. In pa…
Distance metric learning is a successful way to enhance the performance of the nearest neighbor classifier. In most cases, however, the distribution of data does not obey a regular form and may change in different parts of the feature space. Regarding that, this paper proposes a novel local distance metric learning met…
A new robust time series distance metric for k-NN classification.
MultiDendrograms is a Java-written application that computes agglomerative hierarchical clusterings of data. Starting from a distances (or weights) matrix, MultiDendrograms is able to calculate its dendrograms using the most common agglomerative hierarchical clustering methods. The application implements a variable-gro…
Unlike the case of surfaces of topologically finite type, there are several different Teichmüller spaces that are associated to a surface of topological infinite type. These Teichmüller spaces first depend (set-theoretically) on whether we work in the hyperbolic category or in the conformal category. They also depend, …
Paper proposes a new Wasserstein distance for mixtures of radially contoured distributions.