The paper analyzes the generalization of deep neural networks for metric and similarity learning.
problem Lack of rigorous understanding of generalization performance in metric and similarity learning.
method Derive explicit form of true metric, construct structured deep ReLU neural network, establish excess risk bounds.
result Explicit excess risk bounds for metric and similarity learning are derived.
Study evaluates relevance metrics for similarity-based model explanations.
problem Providing understandable explanations for complex model predictions.
method Evaluated three relevance metrics using three tests.
result Cosine similarity of gradients performs best for explanations.
Modified cosine distance improves similarity performance in data with variance and correlation.
problem Limitations of traditional cosine similarity in random variable spaces with variance and correlation.
method Proposed a variance-adjusted cosine distance metric to overcome limitations of traditional cosine similarity.
result Modified cosine distance shows 100% test accuracy in KNN model on the Wisconsin Breast Cancer Dataset.
Deconfounds neural network representation similarity metrics to improve consistency and accuracy.
problem Confounding by population structure in similarity metrics like RSA and CKA.
method Covariate adjustment regression to adjust for confounders.
result Improves detection of semantically similar neural networks and consistency in transfer learning.
CatSIM measures image similarity robustly to small changes.
problem Measuring similarity between images, especially with small perturbations.
method Uses structural similarity image quality paradigm, robust to small location changes.
result Structural similarity between images rated higher when not entirely overlapping.
As a highlighting research topic in the multimedia area, cross-media retrieval aims to capture the complex correlations among multiple media types. Learning better shared representation and distance metric for multimedia data is important to boost the cross-media retrieval. Motivated by the strong ability of deep neura…
We introduce GSimCNN (Graph Similarity Computation via Convolutional Neural Networks) for predicting the similarity score between two graphs. As the core operation of graph similarity search, pairwise graph similarity computation is a challenging problem due to the NP-hard nature of computing many graph distance/simila…
Recently, metric learning and similarity learning have attracted a large amount of interest. Many models and optimisation algorithms have been proposed. However, there is relatively little work on the generalization analysis of such methods. In this paper, we derive novel generalization bounds of metric and similarity …
Algorithm learns similarity metrics for individual fairness.
problem Difficulty in learning similarity metrics for individual fairness.
method Gradient descent and Bradley-Terry model for pairwise comparisons.
result Algorithm converges to ground truth metric for individual fairness.
Similarity found in metrics on special Lie groups.
problem Comparing Riemannian metrics on specific Lie groups.
method Proved all metrics are roughly similar via identity.
result All left-invariant Riemannian metrics are roughly similar.
Paper proposes a supervised similarity framework for corporate bonds using RF proximities.
problem Challenges in measuring similarity for corporate bonds due to noisy data and lack of ground truth.
method Proposes a supervised similarity framework using Random Forest for corporate bonds, introducing a novel metric to evaluate similarities.
result Random Forest outperforms other methods in evaluating similarities for corporate bonds.
We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically distinguishability and symmetry, so that similarity between data points of arbi…
Paper improves image retrieval quality using nonlinear rank approximations.
problem Improving image retrieval quality in high-dimensional feature spaces.
method Computes normalized approximated ranks, converts to similarities, and uses them in a new loss function.
result Significant improvement in image retrieval quality on multiple datasets.
New method approximates Individual Fairness using human judgments.
problem Enforcing fairness in classification tasks.
method Approximates a metric for Individual Fairness based on human queries.
result Constructs hypotheses for metric approximations that generalize.
Similarity metrics are a core component of many information retrieval and machine learning systems. In this work we propose a method capable of learning a similarity metric from data equipped with a binary relation. By considering only the similarity constraints, and initially ignoring the features, we are able to lear…
Paper develops a new similarity metric for predicting stock market returns.
problem Predicting stock returns is challenging due to market stochasticity and various influencing factors.
method Case-based reasoning approach using historical pricing data and a novel similarity metric.
result Demonstrates the benefits of the novel similarity metric in predicting stock market returns.
A new method matches similar regions in non-rigid shapes using spectra of differential operators.
problem Evaluating similarity of non-rigid shapes with partiality.
method Alignment of spectra of differential operators (SI-LBO and regular LBO) on a manifold with multiple metrics.
result Matching spectra outperforms competing methods on standard benchmarks.
Paper tackles multi-label learning by improving SVR for positive semidefinite metrics.
problem Learning positive semidefinite metrics for multi-label and label distribution learning.
method Proposes two methods to overcome SVR's limitation in learning positive semidefinite metrics.
result Demonstrates new methods achieve favorable performance in multi-label and label distribution learning.
This paper formalizes state similarity metrics for reinforcement learning.
problem Leveraging state similarity for reinforcement learning in continuous-state systems.
method Introducing a unified formalism for defining topologies through metrics.
result Established a hierarchy of metrics and demonstrated their theoretical implications.
The crucial importance of metrics in machine learning algorithms has led to an increasing interest in optimizing distance and similarity functions, an area of research known as metric learning. When data consist of feature vectors, a large body of work has focused on learning a Mahalanobis distance. Less work has been …
Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…
Study reveals attention mechanism's similarity computation parallels traditional machine learning.
problem Understanding the essence and principles of attention mechanism in deep learning.
method Examined classic metrics and vector space properties in manifold learning, clustering, and supervised learning to identify key characteristics of similarity computation and information propagation.
result Self-attention mechanism in deep learning adheres to the same principles but operates more flexibly and adaptively.
A CBR system investigates the TS-SS metric for document similarity.
problem Finding the most similar documents for user queries.
method Case-Based Reasoning (CBR) with TS-SS, Euclidean, and Cosine similarity measures.
result TS-SS metric shows surprising inappropriateness for high-dimensional features.
ICE proposes a new loss function for deep metric learning.
problem Deep metric learning from instance-level matching distribution.
method Instance Cross Entropy (ICE) loss function.
result ICE outperforms existing methods on real-world benchmarks.
This research proposes a new distance metric using Isolation Forests.
problem Approximating spatial distance between data points.
method Isolation Forests for outlier detection, transforming separation depth into a distance metric.
result The method produces a distance metric invariant to variable scales and capable of handling non-linear relationships.
The paper proposes tree-based methods for automatically learning similarity measures.
problem Automatically learning similarity measures in feature spaces.
method Formulates similarity learning as a pairwise bipartite ranking problem and uses recursive tree-based ROC optimization.
result Validates iterative partitioning procedures for similarity learning and proposes efficient algorithms.
Proposes a new method to learn distance metrics for semi-supervised learning.
problem Inconsistency between perturbed input sets and lack of pairwise relationship information.
method Metric Learning by Similarity Network (MLSN) co-training with a classification network to learn distance metrics adaptively.
result Performs better than state-of-the-art methods on empirical tasks.
We propose a new method for local distance metric learning based on sample similarity as side information. These local metrics, which utilize conical combinations of metric weight matrices, are learned from the pooled spatial characteristics of the data, as well as the similarity profiles between the pairs of samples, …
New method improves reinforcement learning generalization.
problem Few environments lead to poor generalization in reinforcement learning.
method Integrates sequential structure into representation learning, using a policy similarity metric (PSM) and contrastive embeddings (PSEs).
result PSEs improve generalization across various benchmarks.
Geometric stability measures neural network robustness, distinguishing from similarity metrics.
problem Lack of robustness in neural network representations.
method Introduces geometric stability, quantified by Shesha metric measuring self-consistency.
result Stability and similarity are uncorrelated, revealing distinct properties of neural network robustness.
Improved VAEs learn flat latent spaces for better data similarity.
problem Measuring data similarity in latent spaces using Euclidean metric.
method Extend VAEs to learn flat latent manifolds using Riemannian geometry and regularisation.
result Improved performance on video-tracking benchmarks, nears supervised methods.
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
problem Measuring place function similarity at fine spatial granularity.
method Trajectory embedding to reduce dimensions and measure similarity of place functions.
result Embedding similarity can be a metric proxy for place functions at fine spatial granularity.
A novel criterion selects optimal distance metrics for cell profile analysis.
problem Determining the most accurate distance metric for high-dimensional cell profiles.
method Generalized proposition and corollaries to evaluate and select distance metrics.
result Wasserstein and cosine similarity metrics are optimal for general cases.
Novel method decorrelates batches of triplets for active metric learning.
problem Correlation among triplets degrades active learning performance.
method Proposes a novel method to decorrelate batches of triplets, balancing informativeness and diversity.
result Method outperforms state-of-the-art in active metric learning.
QCML improves bond similarity learning in illiquid markets.
problem Improving similarity learning for illiquid corporate bonds.
method Quantum Cognition Machine Learning (QCML) for supervised distance metric learning.
result QCML outperforms classical tree-based models in high-yield markets.
We argue that robustness of explanations---i.e., that similar inputs should give rise to similar explanations---is a key desideratum for interpretability. We introduce metrics to quantify robustness and demonstrate that current methods do not perform well according to these metrics. Finally, we propose ways that robust…
New method learns psychological similarity spaces for unseen stimuli.
problem Generalizing psychological similarity spaces to new stimuli.
method Learn mapping from raw stimuli to similarity space using ANNs.
result ANNs can successfully map raw stimuli into similarity spaces.
In this paper, we present a novel two-stage metric learning algorithm. We first map each learning instance to a probability distribution by computing its similarities to a set of fixed anchor points. Then, we define the distance in the input data space as the Fisher information distance on the associated statistical ma…
STRAPSim measures ETF portfolio similarity better than existing methods.
problem Measuring portfolio similarity for ETFs and portfolios.
method Semantic, two-level, residual-aware portfolio similarity computation.
result STRAPSim outperforms existing methods in predictive accuracy and ranking alignment.
The paper introduces two new metrics on outer space and shows fixed points for their actions.
problem Analyzing metrics on outer space and their geometric group theory implications.
method Defined and analyzed entropy and pressure metrics on outer space, comparing to Weil-Petersson metric.
result For rank r≥4, the metrics have fixed points in their actions on outer space. A carpet is a metric space homeomorphic to the Sierpinski carpet. We characterize, within a certain class of examples, non-self-similar carpets supporting curve families of nontrivial modulus and supporting Poincaré inequalities. Our results yield new examples of compact doubling metric measure spaces supporting Poinca…
DNN-based cross-modal retrieval has become a research hotspot, by which users can search results across various modalities like image and text. However, existing methods mainly focus on the pairwise correlation and reconstruction error of labeled data. They ignore the semantically similar and dissimilar constraints bet…
New metric captures individual neuron tuning across neural networks.
problem Need a metric that respects individual neuron tuning across different neural networks.
method Derived a 'soft' permutation-based metric using optimal transport theory.
result Metric avoids counter-intuitive outcomes and captures geometric insights.
Large language models learn company embeddings from SEC filings.
problem Lack of a rigorous definition of company similarity.
method Pre-trained and finetuned large language models (LLMs) to learn embeddings from SEC filings.
result LLMs can reproduce GICS classifications and indicate similar financial performance.
Many radiological studies can reveal the presence of several co-existing abnormalities, each one represented by a distinct visual pattern. In this article we address the problem of learning a distance metric for plain radiographs that captures a notion of "radiological similarity": two chest radiographs are considered …
One of the most fundamental problems in machine learning is to compare examples: Given a pair of objects we want to return a value which indicates degree of (dis)similarity. Similarity is often task specific, and pre-defined distances can perform poorly, leading to work in metric learning. However, being able to learn …
Study noncollapsed F-limit metric solitons, proving properties similar to smooth Ricci shrinkers.
problem Understanding noncollapsed F-limit metric solitons in Ricci flow.
method Systematic study and proving properties similar to smooth Ricci shrinkers.
result Proves quadratic lower bound for scalar curvature, local gap theorem, global Sobolev inequality, and optimal volume growth lower bound.
Similarity/Distance measures play a key role in many machine learning, pattern recognition, and data mining algorithms, which leads to the emergence of metric learning field. Many metric learning algorithms learn a global distance function from data that satisfy the constraints of the problem. However, in many real-wor…