Completes the space of vector-valued one-forms on manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method clusters categorical data by learning their optimal order and distance.
Neural networks learn distance-based representations, not just intensity.
CADM proposes a cluster-specific distance metric for categorical data clustering.
CPML efficiently learns new metrics for categorical data.
We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize some of the popular metrics, such as the Jaccard and bag distances on sets, Man…
In the present paper we calculate the Gromov-Hausdorff distance between an arbitrary simplex (a metric space all whose non-zero distances are the same) and a finite metric space whose non-zero distances take two distinct values (so-called -distance spaces). As a corollary, a complete solution to generalized Borsuk p…
Proves compactness for timed-metric spaces using new distance and maps.
Study stability of curvature-dimension condition for negative dimensions.
This work briefly explores the possibility of approximating spatial distance (alternatively, similarity) between data points using the Isolation Forest method envisioned for outlier detection. The logic is similar to that of isolation: the more similar or closer two points are, the more random splits it will take to se…
Sharp bounds for distortion risk metrics under uncertain distributions.
Nearest Neighbors Algorithm is a Lazy Learning Algorithm, in which the algorithm tries to approximate the predictions with the help of similar existing vectors in the training dataset. The predictions made by the K-Nearest Neighbors algorithm is based on averaging the target values of the spatial neighbors. The selecti…
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are messy and values are noisy. The distances between the data points are …
In light of the power problems of statistical tests and undisciplined use of alpha-based statistics to compare models, this paper proposes a unified set of distance-based performance metrics, derived as the square root of the sum of squared alphas and squared standard errors. The Bayesian investor views model performan…
-NN classifier is one of the most famous classification algorithms, whose performance is crucially dependent on the distance metric. When we consider the distance metric as a parameter of -NN, learning an appropriate distance metric for -NN can be seen as minimizing the empirical risk of -NN. In this paper,…
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from \emph{unknown} continuous distributions, which form clusters with each cluster containing a composite set of closely located distributions (based on a certain distance metric between di…
Estimates path-valued data using signature metrics and local kernels.
This paper focuses on the study of open curves in a Riemannian manifold M, and proposes a reparametrization invariant metric on the space of such paths. We use the square root velocity function (SRVF) introduced by Srivastava et al. to define a Riemannian metric on the space of immersions M'=Imm([0,1],M) by pullback of…
Study shows effective resistance distance yields more accurate network barycenter than Hamming distance.
Large scale agglomerative clustering is hindered by computational burdens. We propose a novel scheme where exact inter-instance distance calculation is replaced by the Hamming distance between Kernelized Locality-Sensitive Hashing (KLSH) hashed values. This results in a method that drastically decreases computation tim…
New method improves Wasserstein distance for large-scale data.
We generalize Mallows model to learn distance metrics from data.
PolyGraph Discrepancy improves graph generative model evaluation.
A new distance metric derived from information theory and estimation theory.
Study finite-energy metrics over complex manifold degenerations.
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying geometry between outcomes. The value of being sensitive to this geometry has been demo…
In this work we explore the use of metric index structures, which accelerate nearest neighbor queries, in the scenario where we need to interleave insertions and queries during deployment. This use-case is inspired by a real-life need in malware analysis triage, and is surprisingly understudied. Existing literature ten…
We study the action of the elements of the mapping class group of a surface of finite type on the Teichmüller space of that surface equipped with Thurston's asymmetric metric. We classify such actions as elliptic, parabolic, hyperbolic and pseudo-hyperbolic, depending on whether the translation distance of such an elem…
Defines a measure of knot concordance using cobordism distance.
The study connects lamination and orbit closures in hyperbolic manifolds.
We bound the value of the Casson invariant of any integral homology 3-sphere by a constant times the distance-squared to the identity, measured in any word metric on the Torelli group $\T$, of the element of $\T$ associated to any Heegaard splitting of . We construct examples which show this bound is asymptotica…
A new Riemannian metric on curve spaces is complete and smooth.
Graph neural network learns graph distances effectively.
Complex embeddings handle non-metric proximity data better than traditional methods.
We introduce a new metric to evaluate corruption robustness of ML classifiers.
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transf…
The paper develops formulas for hyperbolic simplices based on edge lengths.
A new numerical framework simplifies elastic surface matching and comparison.
Kernel regression is a popular non-parametric fitting technique. It aims at learning a function which estimates the targets for test inputs as precise as possible. Generally, the function value for a test input is estimated by a weighted average of the surrounding training examples. The weights are typically computed b…
Paper proposes a chi-square test for distance correlation.
Paper calculates distances between strata in Teichmüller space, proving a constant separation.
Extends manifold learning to non-Euclidean metrics.
FOCAL tackles offline meta-reinforcement learning with efficient task inference and behavior regularization.
Distance metric learning is an important component for many tasks, such as statistical classification and content-based image retrieval. Existing approaches for learning distance metrics from pairwise constraints typically suffer from two major problems. First, most algorithms only offer point estimation of the distanc…
New metric learning approach for tree data reduces computation cost.
A new robust time series distance metric for k-NN classification.
Study robustness of polynomial neural networks using algebraic geometry.