Topolow embeds dissimilarity data into Euclidean space robustly against non-metricity and sparsity.
problem Embedding dissimilarity data into Euclidean space when dissimilarities are non-metric or sparse.
method Topolow uses a physics-inspired, gradient-free optimization framework to maximize likelihood under a Laplace error model.
result Topolow outperforms standard MDS methods in reconstructing sparse and non-Euclidean data.
The paper proposes scalable methods for selecting prototypes from large dissimilarity datasets.
problem Selecting good prototypes from large dissimilarity datasets.
method Genetic algorithms, dissimilarity-based hashing, unsupervised and supervised criteria.
result The methods select good prototypes efficiently from large datasets.
CDM improves fingerprinting-based positioning accuracy.
problem Quantifying similarity of collections with missing attributes.
method Combining vector-based distance metrics and set operations.
result Improves positioning accuracy by 5% on average.
Paper tackles multi-label learning by improving SVR for positive semidefinite metrics.
problem Learning positive semidefinite metrics for multi-label and label distribution learning.
method Proposes two methods to overcome SVR's limitation in learning positive semidefinite metrics.
result Demonstrates new methods achieve favorable performance in multi-label and label distribution learning.
New method generates adversarial images under various non-smooth metrics.
problem Adversarial perturbations misclassify deep neural networks.
method Proposes an attack methodology for non-ℓp adversarial dissimilarity metrics. result ProxLogBarrier outperforms existing methods and reveals new perturbation types.
We investigate metric learning in the context of dynamic time warping (DTW), the by far most popular dissimilarity measure used for the comparison and analysis of motion capture data. While metric learning enables a problem-adapted representation of data, the majority of methods has been proposed for vectorial data onl…
New RF dissimilarity measures improve multi-view learning accuracy.
problem Improving multi-view learning accuracy in HDLSS problems.
method Modified Random Forest proximity measures for HDLSS multi-view classification.
result Second method significantly more accurate than other state-of-the-art methods.
We consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. Similarity-based anomaly detection algorithms detect abnormally large amounts of similarity or dissimilarity, e.g.~as measured by nearest neighbor Euclidean distances between a test sam…
CLARITY compares dissimilar datasets, identifying structural and relationship inconsistencies.
problem Integrating qualitatively different datasets from various disciplines.
method Non-parametric approach decomposing similarities into structural and relationship components.
result Identifies and interprets inconsistencies between datasets.
Paper introduces k-DTW for robust curve comparison.
problem Robust dissimilarity measure for polygonal curves.
method Introduces k-Dynamic Time Warping (k-DTW) as a novel dissimilarity measure.
result k-DTW is more robust to outliers and has stronger metric properties than DTW.
Hybrid clustering combines partitional and hierarchical clustering for computational effectiveness and versatility in cluster shape. In such clustering, a dissimilarity measure plays a crucial role in the hierarchical merging. The dissimilarity measure has great impact on the final clustering, and data-independent prop…
Metric learning enhances combinatorial coverage metrics' ability to predict classification errors.
problem Dataset dependence of combinatorial coverage metrics in anticipating classification errors.
method Metric learning to improve latent space separation of data classes.
result Metric learning increases SDCCMs' ability to distinguish between correctly and incorrectly classified data.
Paper develops a new objective for hierarchical clustering in Euclidean space.
problem Hierarchical clustering in Euclidean space with dissimilarity scores.
method Develops a new global objective and connects it to bisecting k-means.
result Optimal 2-means solution approximates the new objective, proving bisecting k-means optimizes a natural global objective.
DMAE uses neural networks to cluster data with flexible dissimilarity functions.
problem Clustering data with complex dissimilarity functions.
method Integrates a dissimilarity mixture model into deep learning architectures.
result DMAE achieves competitive clustering accuracy compared to other methods.
Defines metrics to compare neural network representations.
problem Comparing neural network representations across different architectures and tasks.
method Developed a family of metric spaces and modified existing measures to quantify representational dissimilarity.
result Identified relationships between neural representations and anatomical features.
As a highlighting research topic in the multimedia area, cross-media retrieval aims to capture the complex correlations among multiple media types. Learning better shared representation and distance metric for multimedia data is important to boost the cross-media retrieval. Motivated by the strong ability of deep neura…
Random Forest proximity measures for multi-view classification.
problem Combining multiple heterogeneous data views for classification.
method Building dissimilarity representations for each view, fusing them dynamically.
result Dynamic View Selection improves multi-view classification performance.
New method clusters stationary stochastic processes using covariance-based dissimilarity.
problem Clustering wide-sense stationary ergodic stochastic processes.
method Covariance-based dissimilarity measure with consistent algorithms for offline and online clustering.
result Asymptotically consistent algorithms for efficient clustering.
Advances robustness of metric learning by adversarial margin in input space.
problem Improving robustness of metric learning algorithms.
method Imposing adversarial margin in input space, minimizing perturbation loss.
result Enlarged adversarial margin improves generalization and robustness.
DNN-based cross-modal retrieval has become a research hotspot, by which users can search results across various modalities like image and text. However, existing methods mainly focus on the pairwise correlation and reconstruction error of labeled data. They ignore the semantically similar and dissimilar constraints bet…
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
New method uses cohomology to quantify molecular similarity.
problem Quantifying structural dissimilarity in molecular data.
method Gromov-Hausdorff ultrametric based on simplicial complexes and cohomology.
result Demonstrates effectiveness in clustering organic-inorganic halide perovskite structures.
SQFA learns features maximizing Fisher-Rao distance for better classification.
problem Improving classification accuracy through feature learning.
method SQFA learns linear features maximizing Fisher-Rao distance between class-conditional distributions.
result SQFA-H features achieve the best classification accuracy.
New framework to test neural network representation similarity measures.
problem Disagreements among dissimilarity measures in neural networks.
method Statistical testing framework to evaluate measures based on functional behavior.
result Current metrics have different weaknesses; a classical baseline performs surprisingly well.
A new metric learning scheme for structured data combining graph and feature-space information.
problem Learning a metric from structured data while respecting metric constraints.
method Training metric-constrained linear combinations of dissimilarity matrices, applying graph-based optimization under constraints.
result Our approach can reduce computational complexity by one order of magnitude for some cases.
We propose a framework, named Aggregated Wasserstein, for computing a dissimilarity measure or distance between two Hidden Markov Models with state conditional distributions being Gaussian. For such HMMs, the marginal distribution at any time spot follows a Gaussian mixture distribution, a fact exploited to softly matc…
New method learns psychological similarity spaces for unseen stimuli.
problem Generalizing psychological similarity spaces to new stimuli.
method Learn mapping from raw stimuli to similarity space using ANNs.
result ANNs can successfully map raw stimuli into similarity spaces.
This paper considers networks where relationships between nodes are represented by directed dissimilarities. The goal is to study methods that, based on the dissimilarity structure, output hierarchical clusters, i.e., a family of nested partitions indexed by a connectivity parameter. Our construction of hierarchical cl…
DID measures similarity invariant to diffeomorphisms.
problem Measuring similarity invariance to diffeomorphisms.
method DID measures similarity as the solution to an optimization problem in a Reproducing Kernel Hilbert Space.
result DID is invariant to diffeomorphisms and can be efficiently approximated.
New metric solves correspondence problem for robotic arm imitation learning.
problem Establishing corresponding states and actions between different robotic arms.
method Introducing a distance measure between dissimilar robotic arms and using it as a loss function.
result The distance measure effectively learns imitation policies by minimizing distance between robotic arms.
Paper introduces a new method for learning with distributions using dissimilarity measures.
problem Learning with probability distributions using dissimilarity measures.
method Introduces embeddings based on dissimilarity of distributions to templates, extending similarity theory to population distributions.
result Proves that dissimilarity theory holds for empirical distributions and shows better performance of Wasserstein distance embedding.
New dissimilarity measures enhance affinity propagation for complex network clustering.
problem Improving community detection in complex networks using affinity propagation.
method Leverage network latent geometry to design dissimilarity matrices.
result Affinity propagation outperforms state-of-the-art methods in community detection.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.
A method for classification using pairwise similarities and unlabeled data.
problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.
Neuc-MDS extends MDS for non-Euclidean data.
problem Limitations of classical MDS with non-Euclidean data.
method Generalizes inner product to symmetric bilinear forms, optimizes eigenvalues of dissimilarity Gram matrix.
result Optimizes STRESS for non-Euclidean data.
FPI methods compute barycenters of Gaussian sets for various dissimilarity measures.
problem Efficiently compute barycenters of Gaussian sets for multiple dissimilarity measures.
method Fixed-Point Iterations (FPI) for several dissimilarity measures.
result FPI provides a useful toolbox for fusion/reduction of Gaussian sets.
Paper proposes a new efficient transport-based dissimilarity measure for time series classification.
problem Classifying time series with warping distortions.
method Defining a problem statement, proposing an Optimal Transport-based dissimilarity measure.
result The proposed method can solve the time series classification problem with reduced computational cost.
ClustGeo uses Ward-like clustering with spatial constraints in R.
problem Hierarchical clustering with spatial/geographical constraints.
method Ward-like hierarchical clustering algorithm with two dissimilarity matrices and a mixing parameter.
result Determines optimal spatial contiguity without sacrificing variable quality.
New model for learning from noisy human comparisons, improving search efficiency.
problem Designing efficient algorithms for content search with noisy human feedback.
method Introducing a weak oracle model for comparison-based queries and developing WORCS-I and WORCS-II algorithms.
result Provable algorithms locating target objects with close to entropy of target distribution.
New homology theory for graphs detects subdivisions and homology manifolds.
problem Defining a dissimilarity metric for graphs.
method Filtration on simplicial homology, using bi-colourings of vertices.
result The überhomology vanishes in lowest degree for subdivisions and coincides with fundamental class for homology manifolds.
Locality sensitive hashing (LSH) is a powerful tool for sublinear-time approximate nearest neighbor search, and a variety of hashing schemes have been proposed for different dissimilarity measures. However, hash codes significantly depend on the dissimilarity, which prohibits users from adjusting the dissimilarity at q…
A new metric for comparing HMMs, especially GMM-HMMs, without Monte Carlo samples.
problem Comparing Hidden Markov Models (HMMs) with Gaussian conditional distributions.
method Aggregated Wasserstein metric based on optimal transport between Gaussian mixtures.
result The Aggregated Wasserstein metric is a semi-metric that can be computed efficiently and is invariant to state relabeling.
Paper introduces a method for supervised hierarchical clustering with Exponential Linkage.
problem Discrepancy between training and clustering objectives in supervised clustering.
method Tightly couples supervised training of dissimilarity function with hierarchical clustering, using Exponential Linkage.
result Joint training procedure consistently matches or outperforms other methods, improving dendrogram purity by up to 8 points.
Method transfers feature representation from large to small models using perception coherence.
problem Transfer feature representation from large to small models.
method Defines perception coherence, proposes loss function to minimize.
result Method outperforms or achieves on-par performance compared to strong baseline methods.
The one-class classification problem is a well-known research endeavor in pattern recognition. The problem is also known under different names, such as outlier and novelty/anomaly detection. The core of the problem consists in modeling and recognizing patterns belonging only to a so-called target class. All other patte…
We propose a novel method to determine the dissimilarity between subjects for functional data clustering. Spline smoothing or interpolation is common to deal with data of such type. Instead of estimating the best-representing curve for each subject as fixed during clustering, we measure the dissimilarity between subjec…
BERT improved for propaganda detection with imbalanced, dissimilar data.
problem BERT struggles with dissimilar imbalanced datasets in propaganda detection.
method Cost-sensitive BERT with dissimilarity measure for imbalanced, dissimilar datasets.
result Achieved second-highest score on sentence-level propaganda classification.
A new method for visualizing group dissimilarity using Likelihood Ratio Test.
problem Explaining group differences when null hypothesis is rejected.
method The Merging Path Plot and factorMerger package.
result Visualization of group dissimilarity using Likelihood Ratio Test.