Metric learning enhances combinatorial coverage metrics' ability to predict classification errors.
problem Dataset dependence of combinatorial coverage metrics in anticipating classification errors.
method Metric learning to improve latent space separation of data classes.
result Metric learning increases SDCCMs' ability to distinguish between correctly and incorrectly classified data.
CPML efficiently learns new metrics for categorical data.
problem Metric learning for categorical data.
method CPML (categorical projected metric learning) using Schatten p-norms.
result CPML provides efficient metric learning with improved accuracy.
The need for appropriate ways to measure the distance or similarity between data is ubiquitous in machine learning, pattern recognition and data mining, but handcrafting such good metrics for specific problems is generally difficult. This has led to the emergence of metric learning, which aims at automatically learning…
Compressing data helps learn Mahalanobis metrics effectively.
problem Learning Mahalanobis metrics in high-dimensional spaces.
method Randomly compress data to train a full-rank metric in a reduced feature space.
result Theoretical guarantees on error for Mahalanobis metric learning, independent of ambient dimension.
A neural network approach to compute stable metrics for numerical simulation data.
problem Computing stable and generalizing metrics for diverse numerical simulation data.
method A Siamese neural network architecture with a specialized loss function trained on a controlled data generation setup.
result LSiM outperforms existing metrics for vector spaces and image-based metrics.
This work evaluates and benchmarks calibration metrics for data-driven regression models.
problem Conflicting results from different calibration metrics make it hard to compare and interpret model performance.
method Systematically extracted and benchmarked 14 regression calibration metrics across various data types and recalibration methods.
result Many metrics disagree on the same recalibration result, highlighting the need for careful metric selection.
Estimates mass of static vacuum metrics with small Bartnik data.
problem Estimating mass of static vacuum metrics with small perturbations.
method Second-order mass estimation using Bartnik data.
result New upper bound on Bartnik mass to fifth order.
New metric learning approach for tree data reduces computation cost.
problem Efficiently computing distances between ordered labeled trees.
method Introduced pq-grams and a differentiable weighted pq-gram distance, combined with LMNN for optimization.
result Significantly reduces computation time for tree classification problems.
CADM proposes a cluster-specific distance metric for categorical data clustering.
problem Inadequate distance metrics for categorical data, especially varying within clusters.
method Cluster-customized adaptive distance metric for categorical data.
result Achieved competitive performance in categorical data clustering.
Riemannian metric learning improves data representation across various fields.
problem Traditional distance metrics fail to capture intrinsic data geometry.
method Leverages differential geometry to model data on Riemannian manifolds.
result Demonstrates remarkable success in diverse domains.
Paper introduces metrics to evaluate missing data imputation without ground truth.
problem Handling missing data in time series without ground truth.
method Introduces Wasserstein distance (WD) and Jensen-Shannon divergence (JSD) as metrics to evaluate imputation quality.
result WD and JSD are effective metrics for assessing missing data imputation quality.
This paper explores the impact of metric choice on Fréchet regression.
problem Choosing the right metric for Fréchet regression in complex data.
method Review and extensive numerical studies of existing dimension reduction methods.
result Different metrics significantly affect the estimation of central and central mean space.
Metric learning makes it plausible to learn distances for complex distributions of data from labeled data. However, to date, most metric learning methods are based on a single Mahalanobis metric, which cannot handle heterogeneous data well. Those that learn multiple metrics throughout the space have demonstrated superi…
The development of a metric for structural data is a long-term problem in pattern recognition and machine learning. In this paper, we develop a general metric for comparing nonlinear dynamical systems that is defined with Perron-Frobenius operators in reproducing kernel Hilbert spaces. Our metric includes the existing …
New metric learning algorithm for imbalanced data.
problem Learning metrics in imbalanced datasets.
method Designing a new Mahalanobis metric learning algorithm (IML) for class imbalance.
result IML shows efficiency in learning metrics for imbalanced data.
We propose a low-rank approach to learning a Mahalanobis metric from data. Inspired by the recent geometric mean metric learning (GMML) algorithm, we propose a low-rank variant of the algorithm. This allows to jointly learn a low-dimensional subspace where the data reside and the Mahalanobis metric that appropriately f…
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are messy and values are noisy. The distances between the data points are …
Most of metric learning approaches are dedicated to be applied on data described by feature vectors, with some notable exceptions such as times series, trees or graphs. The objective of this paper is to propose a metric learning algorithm that specifically considers relational data. The proposed approach can take benef…
Proves existence of static vacuum metrics with specific boundary data.
problem Existence of static vacuum metrics with prescribed boundary data.
method Proves existence and local uniqueness of static vacuum metrics close to the Euclidean metric.
result Existence of static vacuum metrics with prescribed Bartnik boundary data.
Deep metric learning detects anomalies without labels.
problem Unsupervised anomaly detection for high-dimensional data.
method Deep metric learning with end-to-end optimization, data distillation, hard mining.
result Significant performance gains over state-of-the-art methods.
New metrics assess class overlap and imbalance in datasets.
problem Class overlap and imbalance make datasets hard to classify.
method Developed new metrics based on ball coverage by classes.
result Metrics correlate well with classifier performance.
We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize some of the popular metrics, such as the Jaccard and bag distances on sets, Man…
Two simple methods learn fair metrics from data to improve fairness in ML tasks.
problem Lack of widely accepted fair metrics for many ML tasks hinders individual fairness adoption.
method Presented two simple ways to learn fair metrics from various data types.
result Fair training with learned metrics improves fairness on three ML tasks.
Stokes equations help uniquely identify manifold metrics from boundary data.
problem Determining Riemannian metric from boundary Cauchy data.
method Proving uniqueness of metric from Stokes equations Cauchy data.
result Partial derivatives of all orders of the metric on the boundary are uniquely determined.
Distance metric learning (DML) has been studied extensively in the past decades for its superior performance with distance-based algorithms. Most of the existing methods propose to learn a distance metric with pairwise or triplet constraints. However, the number of constraints is quadratic or even cubic in the number o…
Novel method transfers orometric measures to metric data sets, identifying key items.
problem Identifying key items in metric data sets like knowledge graphs.
method Transfers orometric measures to bounded metric spaces, using 'isolation' and 'prominence' functions.
result Identifies structurally relevant items in geographic data sets of Germany and France.
New proofs confirm travel time data determine simple metrics on a disc.
problem Determining a simple Riemannian metric from travel time data.
method Proofs based on Myers-Steenrod theorem, Lipschitz-type stability estimate.
result Travel time data determine a simple Riemannian metric on a disc up to natural gauge.
We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically distinguishability and symmetry, so that similarity between data points of arbi…
A simpler metric for latent space geometry.
problem Complexity in capturing geometric structure of data manifolds.
method Prior-based approximate latent Riemannian metric.
result The proposed metric is simple, efficient, and robust.
CAI automates extraction and validation of corporate GHG emission metrics.
problem Manual extraction of corporate GHG emission metrics is labor-intensive and error-prone.
method CAI uses LLMs to automate extraction and validation of metrics from corporate disclosures.
result CAI improves data collection efficiency and accuracy by automating the process.
This paper improves probabilistic latent models on hyperbolic spaces.
problem Uncertainty in predictions due to geodesics crossing low-data regions.
method Augmenting hyperbolic manifold with a pullback metric for probabilistic pullback metrics.
result Geodesics on pullback metric respect both geometry and data distribution, reducing uncertainty.
A new robust time series distance metric for k-NN classification.
problem Robustness against arbitrary data contamination in time series classification.
method Proposes a novel distance metric with worst-case O(nlogn) complexity. result Demonstrates competitive classification accuracy in k-NN time series classification.
Optimal transport learns Riemannian metrics for evolving probability measures.
problem Learning metrics for evolving probability measures on Riemannian manifolds.
method Neural parametrization of a metric tensor via optimal transport, alternating optimization scheme.
result Improved trajectory inference on scRNA and bird migration data.
Locally adaptive nearest neighbors improve automated systems' performance and are easier to interpret.
problem Improving automated systems' performance and interpretability.
method Developed a method for k nearest neighbors algorithms to learn locally adaptive metrics.
result Locally adaptive metrics improve performance and are interpretable.
A framework evaluates synthetic tabular data quality objectively.
problem Lack of an objective interpretation of tabular data metrics.
method Proposes a single mathematical objective for synthetic tabular data distribution, structurally decomposes it, and unifies existing metrics.
result Synthesizers that represent tabular structure outperform other methods, especially on smaller datasets.
A new metric-based principal curve method learns 1D manifolds from spatial data.
problem Learning 1D manifolds from spatial data.
method Metric-based Principal Curve (MPC) approach.
result The method effectively learns the shape of 1D manifolds from synthetic and real datasets.
Metrics stabilize persistent homology in data analysis.
problem Stabilizing invariants for characterizing connectivity structures in data.
method Using contour functions to define metrics for rank invariants.
result Optimal contours provide robust descriptors of spatial patterns.
New TDA approach using Finsler metrics.
problem Traditional TDA concepts and methods.
method Introducing Finsler metrics for TDA.
result Relevance of Finsler metrics to TDA.
A new metric learning scheme for structured data combining graph and feature-space information.
problem Learning a metric from structured data while respecting metric constraints.
method Training metric-constrained linear combinations of dissimilarity matrices, applying graph-based optimization under constraints.
result Our approach can reduce computational complexity by one order of magnitude for some cases.
Article establishes criteria for multiplier Hermitian-Einstein metrics on KSM-manifolds.
problem Existence of multiplier Hermitian-Einstein metrics on Fano manifolds.
method Criterion based on KSM-data and continuous paths connecting solitons.
result Explicit example of a KSM-manifold with a family of multiplier Hermitian-Einstein metrics.
New method makes quality metrics scale-invariant for high-dimensional data.
problem Scale sensitivity in quality metrics affects the accuracy of data projections.
method Analytical and empirical investigation of stress and KL divergence; introduction of a scale-invariant technique.
result The proposed technique accurately captures expected behavior and makes metrics scale-invariant.
We study fairness in collaborative-filtering recommender systems, which are sensitive to discrimination that exists in historical data. Biased data can lead collaborative-filtering methods to make unfair predictions for users from minority groups. We identify the insufficiency of existing fairness metrics and propose f…
We investigate metric learning in the context of dynamic time warping (DTW), the by far most popular dissimilarity measure used for the comparison and analysis of motion capture data. While metric learning enables a problem-adapted representation of data, the majority of methods has been proposed for vectorial data onl…
Estimates means in metric spaces using quantization.
problem No practical estimator for Fréchet means in all metric spaces.
method Introduced estimators based on random quantization and data-driven partitioning.
result Universal consistency of estimators across separable metric spaces and Banach spaces.
Control Contraction Metrics (CCMs) provide a nonlinear controller design involving an offline search for a Riemannian metric and an online search for a shortest path between the current and desired trajectories. In this paper, we generalize CCMs to Finsler geometry, allowing the use of non-Riemannian metrics. We provid…
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
problem Inconsistent validation of synthetic data generated for small sample sizes.
method Proposes a normalized Bottleneck distance metric to evaluate synthetic tabular data.
result Common metrics like propensity scoring and MMD fail for small datasets, showing instability and high variability.
A new method clusters categorical data by learning their optimal order and distance.
problem Clustering categorical data lacks a well-defined metric space.
method Order distance metric learning for categorical data.
result Superior clustering accuracy on categorical and mixed datasets.
Paper develops metrics for random dynamical systems using vector-valued RKHSs.
problem Creating metrics for random nonlinear dynamical systems.
method Develops metrics on random dynamical systems using Perron-Frobenius operators in vector-valued reproducing kernel Hilbert spaces (vvRKHSs). Uses operator-valued kernels and time-wise independence criteria.
result Extends existing metrics for deterministic systems and introduces kernel maximal mean discrepancy for random processes.