Enhanced ensemble clustering via fast propagation of cluster-wise similarities.
problem Challenges in exploring higher-level granularity and multi-scale indirect relationships in ensemble clustering.
method A novel ensemble clustering approach based on fast propagation of cluster-wise similarities via random walks.
result Proposes a new cluster-wise similarity matrix and consensus functions to achieve enhanced co-association and consensus clustering.
Cluster-wise linear regression (CLR), a clustering problem intertwined with regression, is to find clusters of entities such that the overall sum of squared errors from regressions performed over these clusters is minimized, where each cluster may have different variances. We generalize the CLR problem by allowing each…
In several application domains, high-dimensional observations are collected and then analysed in search for naturally occurring data clusters which might provide further insights about the nature of the problem. In this paper we describe a new approach for partitioning such high-dimensional data. Our assumption is that…
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the L2 Wasserstein distance between dist…
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
problem Clustering and generating synthetic data from heterogeneous tabular datasets with hidden cluster structure.
method Developed MMM and MMMsynth algorithms for clustering and synthetic data generation.
result MMMsynth algorithm outperforms other literature tabular-data generators and approaches real data performance.
This study improves child welfare risk models using clustering methods.
problem Improving predictive risk models for child maltreatment.
method Integration of PCA and K-Means clustering.
result No significant difference in model performance across clusters, but better for younger children.
GSimCNN predicts graph similarity using CNNs, outperforming existing methods.
problem Challenging pairwise graph similarity computation due to NP-hardness.
method Graph Edit Distance (GED) as core metric, GSimCNN (Convolutional Neural Networks).
result State-of-the-art performance on graph similarity search.
Enhances clustering performance by integrating tensor similarity.
problem Noise contamination and imbalance in samples or features hinder accurate clustering.
method Proposes a high-order similarity matrix from tensor similarity, which captures spatial information and complements pairwise similarity.
result The proposed IPS2 method significantly outperforms previous similarity-based methods on real-world datasets.
Introduces CHL, a new loss function for continuous similarity learning.
problem Binary similarity learning limitations.
method CHL is a novel loss function that generalizes histogram loss to continuous similarities.
result CHL solves a wider range of tasks including similarity learning, representation learning, and data visualization.
Hierarchical clustering based on pairwise similarities is a common tool used in a broad range of scientific applications. However, in many problems it may be expensive to obtain or compute similarities between the items to be clustered. This paper investigates the hierarchical clustering of N items based on a small sub…
Deconfounds neural network representation similarity metrics to improve consistency and accuracy.
problem Confounding by population structure in similarity metrics like RSA and CKA.
method Covariate adjustment regression to adjust for confounders.
result Improves detection of semantically similar neural networks and consistency in transfer learning.
Quantum networks learn task-dependent asymmetric similarity measures.
problem Challenges of conventional distance functions in capturing meaningful similarity.
method GQSim: Quantum networks for learning task-dependent (a)symmetric similarity.
result Quantum similarity measures extract salient features and achieve theoretically guaranteed performance.
Defines a similarity measure for classification distributions.
problem Measuring similarity between classification distributions.
method Proposes task similarity, a novel measure quantifying performance of source distributions on target distributions.
result Empirical task similarity correlates with transfer efficiency and semantic similarity of source distributions.
In this paper, we investigate the similarity transformations in the Minkowski-n space. We study the geometric invariants of non-null curves under the similarity transformations. Besides, we extend the fundamental theorem for a non-null curve according to a similarity motion. We determine all non-null self-similar curve…
Proposes neural similarity for CNNs to enhance flexibility and performance.
problem Limited flexibility of inner product-based convolution in CNNs.
method Introduces neural similarity as a learnable parametric similarity measure, and proposes NSL for adaptive learning from data.
result Dynamic neural similarity improves flexibility and performance in visual recognition and few-shot learning.
Automates similarity measure construction from data.
problem Defining similarity measures analytically is challenging.
method Applies machine learning to learn similarity measures from data.
result Data-driven similarity measure outperforms state of the art methods.
Similarity-based clustering and semi-supervised learning methods separate the data into clusters or classes according to the pairwise similarity between the data, and the pairwise similarity is crucial for their performance. In this paper, we propose a novel discriminative similarity learning framework which learns dis…
Modified cosine distance improves similarity performance in data with variance and correlation.
problem Limitations of traditional cosine similarity in random variable spaces with variance and correlation.
method Proposed a variance-adjusted cosine distance metric to overcome limitations of traditional cosine similarity.
result Modified cosine distance shows 100% test accuracy in KNN model on the Wisconsin Breast Cancer Dataset.
Advocates Tversky's model for image similarity learning.
problem Learning Tversky similarity measures from image data.
method Computational approach using Tversky's ratio model.
result Performs well compared to existing methods on image datasets.
Paper presents a new trie for integer sketches to improve similarity searches.
problem Efficient similarity searches on integer sketches.
method Introduces a novel b-bit sketch trie that leverages succinct data structures. result Significantly improves search time and space-efficiency of similarity searches.
WIPS optimizes inner product weights to approximate various similarities.
problem Learning high-quality node representations and accurate similarities.
method Weighted inner product similarity (WIPS) with adjustable weights.
result WIPS can approximate arbitrary general similarities including positive definite and indefinite kernels.
Method measures weight similarity in neural networks using normalization and statistical inference.
problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.
We prove that the only self-similar surfaces of Euclidean 3-space which are foliated by circles are the self-similar surfaces of revolution discovered by S. Angenent and that the only ruled, self-similar surfaces are the cylinders over planar self-similar curves.
Improves confidence calibration in neural networks by smoothing labels based on class similarity.
problem Improving confidence calibration in deep neural networks for safety-critical applications.
method Proposes a novel label smoothing technique where label values are based on similarities with the reference class, using different similarity measurements.
result Consistently outperforms state-of-the-art calibration techniques on various datasets and network architectures.
CatSIM measures image similarity robustly to small changes.
problem Measuring similarity between images, especially with small perturbations.
method Uses structural similarity image quality paradigm, robust to small location changes.
result Structural similarity between images rated higher when not entirely overlapping.
Proposes DNN-based speaker embedding correlated with subjective inter-speaker similarity for speech synthesis.
problem Inadequate speaker representation for open speakers not in training data.
method Two training algorithms using inter-speaker similarity matrices: similarity vector embedding and similarity matrix embedding.
result Proposed algorithms learn speaker embedding highly correlated with subjective inter-speaker similarity.
This work shows cosine similarity is equivalent to Pearson correlation for word vectors, but not all vectors are suitable for cosine.
problem The use of cosine similarity for semantic textual similarity is often taken for granted, despite its limitations.
method Characterized cases where Pearson correlation is unfit and introduced rank correlation as an alternative.
result Pearson correlation is equivalent to cosine similarity for many word vectors but not all, and rank correlation can improve performance.
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
problem Measuring place function similarity at fine spatial granularity.
method Trajectory embedding to reduce dimensions and measure similarity of place functions.
result Embedding similarity can be a metric proxy for place functions at fine spatial granularity.
New similarity index avoids limitations of CCA in neural networks.
problem Limitations of existing methods in measuring neural network representation similarity.
method Introducing a similarity index based on centered kernel alignment (CKA) to measure representational similarity matrices.
result CKA reliably identifies correspondences between representations in networks trained from different initializations.
STRAPSim measures ETF portfolio similarity better than existing methods.
problem Measuring portfolio similarity for ETFs and portfolios.
method Semantic, two-level, residual-aware portfolio similarity computation.
result STRAPSim outperforms existing methods in predictive accuracy and ranking alignment.
We propose shifted inner-product similarity (SIPS), which is a novel yet very simple extension of the ordinary inner-product similarity (IPS) for neural-network based graph embedding (GE). In contrast to IPS, that is limited to approximating positive-definite (PD) similarities, SIPS goes beyond the limitation by introd…
Exploits class similarity for better machine learning models with confidence labels and projective loss functions.
problem Poor model performance due to confusing similar classes.
method Exploits class similarity with confidence labels and projective loss functions.
result Improved model performance on noisy labels.
BiLRP explains deep similarity models by decomposing scores into feature contributions.
problem Verifying meaningful patterns in complex similarity models.
method Augmenting similarity scores with feature explanations using LRP.
result BiLRP robustly explains deep neural network features and historical document similarities.
The problem of hierarchical clustering items from pairwise similarities is found across various scientific disciplines, from biology to networking. Often, applications of clustering techniques are limited by the cost of obtaining similarities between pairs of items. While prior work has been developed to reconstruct cl…
Lipschitz equivalence of self-similar sets is an important area in the study of fractal geometry. It is known that two dust-like self-similar sets with the same contraction ratios are always Lipschitz equivalent. However, when self-similar sets have touching structures the problem of Lipschitz equivalence becomes much …
New method improves reinforcement learning generalization.
problem Few environments lead to poor generalization in reinforcement learning.
method Integrates sequential structure into representation learning, using a policy similarity metric (PSM) and contrastive embeddings (PSEs).
result PSEs improve generalization across various benchmarks.
New findings clarify the link between distributional closeness and representational similarity.
problem When and why do different neural network representations become similar?
method Identifiability theory, focusing on model families including autoregressive language models.
result Small Kullback-Leibler divergence does not guarantee similar representations.
Neural networks auto-denoise similar inputs, enabling new statistical analysis.
problem Estimating similarity of inputs for neural networks.
method Define and quantify similarity from neural network perspective, using parameter variation impact on outputs.
result Estimate sample density and quantify denoising effect without true labels.
Unified understanding of neural representation similarity measures.
problem Fragmented research landscape of neural network similarity measures.
method Observation and exploration of connections between shape distances and normalized Bures similarity.
result Cosine of the Riemannian shape distance equals normalized Bures similarity.
ROTS improves sentence similarity by incorporating structural information.
problem Measuring sentence similarity with theoretical insights and structural awareness.
method Recursive Optimal Transport (ROT) framework to incorporate structural information.
result ROTS outperforms weakly supervised approaches in sentence similarity tasks.
DBSCAN clustering improved by using nearest neighbour-induced Isolation Similarity.
problem Improving clustering performance of DBSCAN.
method Proposed nearest neighbour method to implement Isolation Similarity.
result DBSCAN clustering performance surpassed by DP algorithm.
New method learns psychological similarity spaces for unseen stimuli.
problem Generalizing psychological similarity spaces to new stimuli.
method Learn mapping from raw stimuli to similarity space using ANNs.
result ANNs can successfully map raw stimuli into similarity spaces.
CLS measures dataset similarity through decision rule performance.
problem Measuring dataset similarity in machine learning, especially for transfer learning and domain adaptation.
method Cross-Learning Score (CLS) measures similarity through bidirectional generalization performance of decision rules, linking to cosine similarity under canonical linear models.
result CLS effectively measures dataset similarity and transferability, validated on synthetic and real-world datasets.
Survey of deep learning methods for graph similarity.
problem Learning a similarity metric among graphs.
method Deep learning models mapping graphs to a target space.
result Systematic taxonomy of existing methods and applications.
New method for estimating firm linkages using CVLs and QCML.
problem Estimating firm linkages for profitable trading strategies.
method Characteristic Vector Linkages (CVLs) and Quantum Cognition Machine Learning (QCML).
result QCML similarity outperforms Euclidean similarity in constructing profitable trading strategies.
New measure quantifies function similarity for optimization.
problem Measuring similarity between functions for optimization.
method Quantifies sub-optimality gaps and operation rules.
result Unified measure for various functional similarities.
Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between them, or a mix of both. Existing methods for relational clustering have strong and …
Method screens similar capsule endoscopic images, reducing doctor workload and improving accuracy.
problem Time-consuming and high error rate in manual inspection of large numbers of similar capsule endoscopic images.
method Structural similarity analysis of visually salient areas and hierarchical clustering.
result 76% reduction in similar images, 100% lesion recall, 18-minute average play time.