Completes the space of vector-valued one-forms on manifolds.
problem Metric incompleteness of the space of full-ranked one-forms.
method Distance equality and quotient structures.
result Concrete description of the metric completion of the space of full-ranked one-forms.
A new method clusters categorical data by learning their optimal order and distance.
problem Clustering categorical data lacks a well-defined metric space.
method Order distance metric learning for categorical data.
result Superior clustering accuracy on categorical and mixed datasets.
Neural networks learn distance-based representations, not just intensity.
problem Understanding how neural networks interpret and learn from internal activations.
method Manipulated ReLU and Absolute Value activations to observe sensitivity to distance and intensity perturbations.
result Neural networks are highly sensitive to small distance-based perturbations, challenging the intensity-based interpretation.
CADM proposes a cluster-specific distance metric for categorical data clustering.
problem Inadequate distance metrics for categorical data, especially varying within clusters.
method Cluster-customized adaptive distance metric for categorical data.
result Achieved competitive performance in categorical data clustering.
CPML efficiently learns new metrics for categorical data.
problem Metric learning for categorical data.
method CPML (categorical projected metric learning) using Schatten p-norms.
result CPML provides efficient metric learning with improved accuracy.
We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize some of the popular metrics, such as the Jaccard and bag distances on sets, Man…
In the present paper we calculate the Gromov-Hausdorff distance between an arbitrary simplex (a metric space all whose non-zero distances are the same) and a finite metric space whose non-zero distances take two distinct values (so-called 2-distance spaces). As a corollary, a complete solution to generalized Borsuk p…
Proves compactness for timed-metric spaces using new distance and maps.
problem Weak convergence of space-times using timed-Hausdorff distance.
method Uses Gromov's original compactness theorem and introduces addresses.
result Establishes compactness theorem for intrinsic timed-Hausdorff convergence.
Study stability of curvature-dimension condition for negative dimensions.
problem Stability of curvature-dimension condition with negative dimension parameters.
method Introduced CD(K, N)-condition for N < 0, defined distance d_{\mathsf{iKRW}}, proved convergence stability.
result Limit structure of converging metric measure spaces remains CD(K, N) for N < 0.
This work briefly explores the possibility of approximating spatial distance (alternatively, similarity) between data points using the Isolation Forest method envisioned for outlier detection. The logic is similar to that of isolation: the more similar or closer two points are, the more random splits it will take to se…
A new method for learning distance metrics for K-NN classification.
problem Improving the performance of K-NN classifier by learning an appropriate distance metric.
method Designing a continuous decision function for K-NN and minimizing its continuous empirical risk function.
result The proposed ANN algorithm outperforms existing methods like LMNN, NCA, and pairwise constraints.
Sharp bounds for distortion risk metrics under uncertain distributions.
problem Modeling risk metrics under distributional uncertainty.
method Established bounds for distortion risk metrics using specific features of underlying distributions.
result Identified worst- and best-case values of distortion risk metrics.
Nearest Neighbors Algorithm is a Lazy Learning Algorithm, in which the algorithm tries to approximate the predictions with the help of similar existing vectors in the training dataset. The predictions made by the K-Nearest Neighbors algorithm is based on averaging the target values of the spatial neighbors. The selecti…
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are messy and values are noisy. The distances between the data points are …
In light of the power problems of statistical tests and undisciplined use of alpha-based statistics to compare models, this paper proposes a unified set of distance-based performance metrics, derived as the square root of the sum of squared alphas and squared standard errors. The Bayesian investor views model performan…
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from \emph{unknown} continuous distributions, which form clusters with each cluster containing a composite set of closely located distributions (based on a certain distance metric between di…
Estimates path-valued data using signature metrics and local kernels.
problem Nonparametric regression and classification for path-valued data.
method Combines signature transform and local kernel regression.
result Establishes convergence bounds and demonstrates competitive accuracy.
This paper focuses on the study of open curves in a Riemannian manifold M, and proposes a reparametrization invariant metric on the space of such paths. We use the square root velocity function (SRVF) introduced by Srivastava et al. to define a Riemannian metric on the space of immersions M'=Imm([0,1],M) by pullback of…
Study shows effective resistance distance yields more accurate network barycenter than Hamming distance.
problem Identifying the best metric for computing the Fréchet mean network.
method Compared the effectiveness of Hamming distance and effective resistance distance in capturing network topology.
result Effective resistance distance produces a more accurate Fréchet mean network.
Large scale agglomerative clustering is hindered by computational burdens. We propose a novel scheme where exact inter-instance distance calculation is replaced by the Hamming distance between Kernelized Locality-Sensitive Hashing (KLSH) hashed values. This results in a method that drastically decreases computation tim…
New method improves Wasserstein distance for large-scale data.
problem High computational cost of Wasserstein distance for large-scale machine learning.
method Augmented Sliced Wasserstein Distances (ASWDs) using neural network mappings.
result ASWDs significantly outperform other Wasserstein variants in synthetic and real-world problems.
We generalize Mallows model to learn distance metrics from data.
problem Learning optimal distance metrics from noisy ranking data.
method Propose Lα distances and develop FPTAS for sampling and MLE. result Strong consistency of estimators for various α and β. PolyGraph Discrepancy improves graph generative model evaluation.
problem Inability of existing metrics to provide an absolute performance measure and comparability across different graph descriptors.
method Approximates Jensen-Shannon distance using binary classifiers trained to distinguish between real and generated graphs.
result PGD provides a more robust and insightful evaluation compared to MMD metrics.
A new distance metric derived from information theory and estimation theory.
problem Developing a robust distance metric for complex signal distributions.
method Information-Estimation Metric (IEM) derived from continuous probability density and denoising errors.
result The IEM is a valid global distance metric that adapts to the geometry of complex distributions.
Study finite-energy metrics over complex manifold degenerations.
problem Finite-energy metrics on complex manifolds with singularities.
method Investigate spaces of plurisubharmonic metrics with finite-energy conditions.
result Complete and geodesic metric structure on finite-energy metrics space.
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
problem Inconsistent validation of synthetic data generated for small sample sizes.
method Proposes a normalized Bottleneck distance metric to evaluate synthetic tabular data.
result Common metrics like propensity scoring and MMD fail for small datasets, showing instability and high variability.
The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying geometry between outcomes. The value of being sensitive to this geometry has been demo…
In this work we explore the use of metric index structures, which accelerate nearest neighbor queries, in the scenario where we need to interleave insertions and queries during deployment. This use-case is inspired by a real-life need in malware analysis triage, and is surprisingly understudied. Existing literature ten…
We study the action of the elements of the mapping class group of a surface of finite type on the Teichmüller space of that surface equipped with Thurston's asymmetric metric. We classify such actions as elliptic, parabolic, hyperbolic and pseudo-hyperbolic, depending on whether the translation distance of such an elem…
Defines a measure of knot concordance using cobordism distance.
problem Measuring how close knots are to being linearly dependent.
method Cobordism distance on cyclic subgroups of knot concordance group.
result Projective space of knot concordance group with integer-valued metric.
The study connects lamination and orbit closures in hyperbolic manifolds.
problem Understanding the geometric and dynamical properties of horocycle orbit closures in Z-covers of compact hyperbolic manifolds. method Exposes connections between distance minimizing laminations and horospherical orbit closures in Z-covers of compact hyperbolic manifolds. Provides novel constructions and explicit descriptions. result Even slight perturbations to hyperbolic metrics can drastically change horocycle orbit closures.
We bound the value of the Casson invariant of any integral homology 3-sphere M by a constant times the distance-squared to the identity, measured in any word metric on the Torelli group $\T$, of the element of $\T$ associated to any Heegaard splitting of M. We construct examples which show this bound is asymptotica…
A new Riemannian metric on curve spaces is complete and smooth.
problem Defining a complete metric on the space of embedded curves.
method Proposed a new Riemannian metric and proved its completeness.
result The proposed metric is complete in multiple senses.
Graph neural network learns graph distances effectively.
problem Maintaining graph distance metric properties.
method GRAPH-BERT based semi-supervised distance metric learning.
result GB-DISTANCE outperforms existing methods.
Complex embeddings handle non-metric proximity data better than traditional methods.
problem Proximities not always metric or inner product-based, causing convergence issues.
method Proposes complex-valued embeddings for non-vectorial data.
result Complex embeddings outperform traditional techniques on benchmarks.
We introduce a new metric to evaluate corruption robustness of ML classifiers.
problem Evaluating corruption robustness of machine learning classifiers.
method We propose a test data augmentation method using minimal class separation distance to derive a robustness distance ε and a metric MSCR.
result The MSCR metric allows interpretable comparison of classifier robustness on different datasets.
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transf…
The paper develops formulas for hyperbolic simplices based on edge lengths.
problem Understanding the geometry of hyperbolic simplices using only edge lengths.
method Develops geometric formulas for hyperbolic simplices based on edge lengths.
result Distance and projection formulas in hyperbolic simplices.
A new numerical framework simplifies elastic surface matching and comparison.
problem Challenging problem in surface comparison and matching in computer vision.
method Relaxing the geodesic boundary constraint using a varifold fidelity metric.
result Flexibility to deal with arbitrary topologies and sampling patterns, scalability to large meshes.
Kernel regression is a popular non-parametric fitting technique. It aims at learning a function which estimates the targets for test inputs as precise as possible. Generally, the function value for a test input is estimated by a weighted average of the surrounding training examples. The weights are typically computed b…
Paper proposes a chi-square test for distance correlation.
problem Testing distance correlation is computationally expensive.
method Proposes a chi-square test for distance correlation, non-parametric, fast, applicable to various metrics.
result Chi-square test exhibits similar power to permutation test and can be valid and universally consistent for testing independence.
Paper calculates distances between strata in Teichmüller space, proving a constant separation.
problem Measuring distances in the Weil-Petersson metric on Teichmüller space.
method Analyzes distances between strata, proving a constant separation and providing bounds.
result Proves the optimal value for minimal separation between strata is a constant δ1,1. Extends manifold learning to non-Euclidean metrics.
problem Applying manifold learning to data in non-Euclidean spaces.
method Generalizes manifold learning to metric spaces and studies conditions for convergence.
result Conditions for the convergence of graph Laplacian in metric spaces.
FOCAL tackles offline meta-reinforcement learning with efficient task inference and behavior regularization.
problem Efficiently adapt RL algorithms to unseen tasks without interactions, addressing bootstrapping errors and robust task inference.
method FOCAL combines behavior regularization, a deterministic context encoder, and a negative-power distance metric for efficient task inference.
result FOCAL outperforms prior algorithms on meta-RL benchmarks, demonstrating computational efficiency.
Distance metric learning is an important component for many tasks, such as statistical classification and content-based image retrieval. Existing approaches for learning distance metrics from pairwise constraints typically suffer from two major problems. First, most algorithms only offer point estimation of the distanc…
New metric learning approach for tree data reduces computation cost.
problem Efficiently computing distances between ordered labeled trees.
method Introduced pq-grams and a differentiable weighted pq-gram distance, combined with LMNN for optimization.
result Significantly reduces computation time for tree classification problems.
A new robust time series distance metric for k-NN classification.
problem Robustness against arbitrary data contamination in time series classification.
method Proposes a novel distance metric with worst-case O(nlogn) complexity. result Demonstrates competitive classification accuracy in k-NN time series classification.
Study robustness of polynomial neural networks using algebraic geometry.
problem Certify robustness radius of polynomial neural networks.
method Metric algebraic geometry, Euclidean distance degree, symbolic elimination, homotopy-continuation methods.
result Found decision boundaries with lower ED degree than generic cubic hypersurfaces.