A new method for visualizing group dissimilarity using Likelihood Ratio Test.
problem Explaining group differences when null hypothesis is rejected.
method The Merging Path Plot and factorMerger package.
result Visualization of group dissimilarity using Likelihood Ratio Test.
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
Locality sensitive hashing (LSH) is a powerful tool for sublinear-time approximate nearest neighbor search, and a variety of hashing schemes have been proposed for different dissimilarity measures. However, hash codes significantly depend on the dissimilarity, which prohibits users from adjusting the dissimilarity at q…
New method assesses data clusterability using ultrametricity.
problem Determining if large datasets are efficiently clusterable.
method Proposes a novel ultrametric-based approach to evaluate clusterability via matrix product.
result Demonstrates the generation of sub-dominant ultrametric from dissimilarity space.
Hybrid clustering combines partitional and hierarchical clustering for computational effectiveness and versatility in cluster shape. In such clustering, a dissimilarity measure plays a crucial role in the hierarchical merging. The dissimilarity measure has great impact on the final clustering, and data-independent prop…
sCSC clusters data without prior assumptions, revealing natural groupings.
problem Clustering data without prior knowledge of its structure.
method sCSC performs binary splittings maximizing dissimilarity, producing a binary tree.
result Clusters emerge naturally from the binary tree, revealing data structure.
DMAE uses neural networks to cluster data with flexible dissimilarity functions.
problem Clustering data with complex dissimilarity functions.
method Integrates a dissimilarity mixture model into deep learning architectures.
result DMAE achieves competitive clustering accuracy compared to other methods.
Random Forest proximity measures for multi-view classification.
problem Combining multiple heterogeneous data views for classification.
method Building dissimilarity representations for each view, fusing them dynamically.
result Dynamic View Selection improves multi-view classification performance.
New method clusters stationary stochastic processes using covariance-based dissimilarity.
problem Clustering wide-sense stationary ergodic stochastic processes.
method Covariance-based dissimilarity measure with consistent algorithms for offline and online clustering.
result Asymptotically consistent algorithms for efficient clustering.
The paper proposes scalable methods for selecting prototypes from large dissimilarity datasets.
problem Selecting good prototypes from large dissimilarity datasets.
method Genetic algorithms, dissimilarity-based hashing, unsupervised and supervised criteria.
result The methods select good prototypes efficiently from large datasets.
This paper considers networks where relationships between nodes are represented by directed dissimilarities. The goal is to study methods that, based on the dissimilarity structure, output hierarchical clusters, i.e., a family of nested partitions indexed by a connectivity parameter. Our construction of hierarchical cl…
DID measures similarity invariant to diffeomorphisms.
problem Measuring similarity invariance to diffeomorphisms.
method DID measures similarity as the solution to an optimization problem in a Reproducing Kernel Hilbert Space.
result DID is invariant to diffeomorphisms and can be efficiently approximated.
Paper introduces a new method for learning with distributions using dissimilarity measures.
problem Learning with probability distributions using dissimilarity measures.
method Introduces embeddings based on dissimilarity of distributions to templates, extending similarity theory to population distributions.
result Proves that dissimilarity theory holds for empirical distributions and shows better performance of Wasserstein distance embedding.
New dissimilarity measures enhance affinity propagation for complex network clustering.
problem Improving community detection in complex networks using affinity propagation.
method Leverage network latent geometry to design dissimilarity matrices.
result Affinity propagation outperforms state-of-the-art methods in community detection.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.
A method for classification using pairwise similarities and unlabeled data.
problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.
FPI methods compute barycenters of Gaussian sets for various dissimilarity measures.
problem Efficiently compute barycenters of Gaussian sets for multiple dissimilarity measures.
method Fixed-Point Iterations (FPI) for several dissimilarity measures.
result FPI provides a useful toolbox for fusion/reduction of Gaussian sets.
Paper proposes a new efficient transport-based dissimilarity measure for time series classification.
problem Classifying time series with warping distortions.
method Defining a problem statement, proposing an Optimal Transport-based dissimilarity measure.
result The proposed method can solve the time series classification problem with reduced computational cost.
ClustGeo uses Ward-like clustering with spatial constraints in R.
problem Hierarchical clustering with spatial/geographical constraints.
method Ward-like hierarchical clustering algorithm with two dissimilarity matrices and a mixing parameter.
result Determines optimal spatial contiguity without sacrificing variable quality.
Finding an informative subset of a large collection of data points or models is at the center of many problems in computer vision, recommender systems, bio/health informatics as well as image and natural language processing. Given pairwise dissimilarities between the elements of a `source set' and a `target set,' we co…
Paper introduces a method for supervised hierarchical clustering with Exponential Linkage.
problem Discrepancy between training and clustering objectives in supervised clustering.
method Tightly couples supervised training of dissimilarity function with hierarchical clustering, using Exponential Linkage.
result Joint training procedure consistently matches or outperforms other methods, improving dendrogram purity by up to 8 points.
The one-class classification problem is a well-known research endeavor in pattern recognition. The problem is also known under different names, such as outlier and novelty/anomaly detection. The core of the problem consists in modeling and recognizing patterns belonging only to a so-called target class. All other patte…
We propose a novel method to determine the dissimilarity between subjects for functional data clustering. Spline smoothing or interpolation is common to deal with data of such type. Instead of estimating the best-representing curve for each subject as fixed during clustering, we measure the dissimilarity between subjec…
BERT improved for propaganda detection with imbalanced, dissimilar data.
problem BERT struggles with dissimilar imbalanced datasets in propaganda detection.
method Cost-sensitive BERT with dissimilarity measure for imbalanced, dissimilar datasets.
result Achieved second-highest score on sentence-level propaganda classification.
New RF dissimilarity measures improve multi-view learning accuracy.
problem Improving multi-view learning accuracy in HDLSS problems.
method Modified Random Forest proximity measures for HDLSS multi-view classification.
result Second method significantly more accurate than other state-of-the-art methods.
Method reveals dissimilarity in alloys' Curie temperatures.
problem Tackles the dissimilarity between rare-earth transition metal binary alloys.
method Ensemble learning with Kernel ridge regression.
result Reveals meaningful relations between alloys' structure and Curie temperature.
Quantitatively assessing relationships between latent variables and observed variables is important for understanding and developing generative models and representation learning. In this paper, we propose latent-observed dissimilarity (LOD) to evaluate the dissimilarity between the probabilistic characteristics of lat…
Multiple instance learning (MIL) is concerned with learning from sets (bags) of objects (instances), where the individual instance labels are ambiguous. In this setting, supervised learning cannot be applied directly. Often, specialized MIL methods learn by making additional assumptions about the relationship of the ba…
We introduce in this paper a new way of optimizing the natural extension of the quantization error using in k-means clustering to dissimilarity data. The proposed method is based on hierarchical clustering analysis combined with multi-level heuristic refinement. The method is computationally efficient and achieves bett…
New method uses cohomology to quantify molecular similarity.
problem Quantifying structural dissimilarity in molecular data.
method Gromov-Hausdorff ultrametric based on simplicial complexes and cohomology.
result Demonstrates effectiveness in clustering organic-inorganic halide perovskite structures.
We consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. Similarity-based anomaly detection algorithms detect abnormally large amounts of similarity or dissimilarity, e.g.~as measured by nearest neighbor Euclidean distances between a test sam…
A new weighted dissimilarity measure reduces positioning errors in feature-based systems.
problem Reducing errors in feature-based positioning systems, especially in areas with high variability.
method Iterative scheme using location-dependent standard deviations as weights.
result Maximum radial positioning error reduced by 40% using the weighted dissimilarity measure.
FedProx algorithm improved for non-smooth and heterogeneous data.
problem Theoretical understanding of FedProx for non-convex federated optimization.
method Local dissimilarity invariant convergence theory through algorithmic stability.
result Convergence guarantees for non-smooth FL problems and minibatch size.
Method identifies regions of maximum dissimilarity in stochastic processes.
problem Comparing local characteristics of two random processes to find periods of maximum dissimilarity.
method Bayesian inference with integrated nested Laplace approximation for stochastic processes.
result Identifies regions of maximum dissimilarity with a certain volume.
In numerous applicative contexts, data are too rich and too complex to be represented by numerical vectors. A general approach to extend machine learning and data mining techniques to such data is to really on a dissimilarity or on a kernel that measures how different or similar two objects are. This approach has been …
Topolow embeds dissimilarity data into Euclidean space robustly against non-metricity and sparsity.
problem Embedding dissimilarity data into Euclidean space when dissimilarities are non-metric or sparse.
method Topolow uses a physics-inspired, gradient-free optimization framework to maximize likelihood under a Laplace error model.
result Topolow outperforms standard MDS methods in reconstructing sparse and non-Euclidean data.
Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented by three distance functions and to identify the optimal distance function for c…
Improved detection of brain tumours in MRIs using latent space dissimilarities.
problem Detecting tumours in brain MRIs using unsupervised learning.
method Slice-wise semi-supervised method based on dissimilarity between latent representations of images and their reconstructions.
result Improved detection results with higher resolution images and better reconstructions.
The Joint Optimization of Fidelity and Commensurability (JOFC) manifold matching methodology embeds an omnibus dissimilarity matrix consisting of multiple dissimilarities on the same set of objects. One approach to this embedding optimizes the preservation of fidelity to each individual dissimilarity matrix together wi…
The paper introduces a clustering method for relational matrices using information value.
problem Clustering relational matrices of similarities or dissimilarities.
method Optimizing the value of information through a deterministic annealing process.
result The method automatically determines the number of clusters without prior specification.
Continuous MDS embeds sequences of dissimilarities in Euclidean space.
problem Embedding sequences of dissimilarities as n increases. method Continuous MDS reformulates MDS for sequences of dissimilarity matrices.
result Uniform convergence of interpolated embeddings.
Unified model for interactive estimation with improved learnability measure.
problem Improving learnability in interactive estimation models.
method Introducing a combinatorial measure (dissimilarity dimension) and a general algorithm with polynomial bounds.
result Unified model subsumes statistical-query learning and structured bandits.
Paper introduces k-DTW for robust curve comparison.
problem Robust dissimilarity measure for polygonal curves.
method Introduces k-Dynamic Time Warping (k-DTW) as a novel dissimilarity measure.
result k-DTW is more robust to outliers and has stronger metric properties than DTW.
This paper considers networks where relationships between nodes are represented by directed dissimilarities. The goal is to study methods for the determination of hierarchical clusters, i.e., a family of nested partitions indexed by a connectivity parameter, induced by the given dissimilarity structures. Our constructi…
Automated protein function prediction is a challenging problem with distinctive features, such as the hierarchical organization of protein functions and the scarcity of annotated proteins for most biological functions. We propose a multitask learning algorithm addressing both issues. Unlike standard multitask algorithm…
CDM improves fingerprinting-based positioning accuracy.
problem Quantifying similarity of collections with missing attributes.
method Combining vector-based distance metrics and set operations.
result Improves positioning accuracy by 5% on average.
We consider general non-Euclidean distance measures between real world objects that need to be classified. It is assumed that objects are represented by distances to other objects only. Conditions for zero-error dissimilarity based classifiers are derived. Additional conditions are given under which the zero-error deci…
Extends RSP model with net flow and capacity constraints for better network analysis.
problem Improving shortest path models with net flows and capacity constraints.
method Developed net flow RSP model and introduced capacity constraints. Proposed algorithms for computing expected routing costs and solving constrained problems using Lagrangian duality.
result Net flow RSP dissimilarity measure is competitive with state-of-the-art dissimilarities.