Optimizes edge coloring in graph bundling for better edge differentiation.
problem Difficulty in identifying origins and destinations of individual edges in strongly bundled graphs.
method Optimizes edge coloring based on pairwise edge strength and origin-destination dissimilarity, solving a nonlinear optimization problem.
result Peacock bundles enhance graph layout comprehensibility with edge differentiation.
A new neural approach for generating origin-destination matrices in ABMs.
problem Challenges in generating origin-destination matrices for ABMs, including discretisation errors and inability to explore multimodal distributions.
method A computationally efficient framework that learns trip intensity through a neural differential equation, operating directly on the discrete combinatorial space.
result Outperforms prior art in terms of reconstruction error and ground truth matrix coverage, at a fraction of the computational cost.
Model predicts passenger origin-destination for online taxi-hailing systems.
problem Predicting passenger origin-destination for efficient transportation planning.
method K-means clustering, non-negative matrix factorization, stacked recurrent neural network.
result Proposed model reduces MAPE by 5-7% for 1-hour windows and 14% for 30-minute windows.
The study introduces measures of collective mobility from aggregated OD data.
problem Understanding large-scale mobility patterns from aggregated data.
method Developed a framework using synthetic and real data to interpret network-level mobility.
result Aggregated mobility measures reveal network structure and flow constraints.
CSTN predicts taxi demand between all regions, overcoming origin-only approaches.
problem Predicting taxi demand between all regions, not just origins.
method Contextualized Spatial-Temporal Network (CSTN) with LSC, TEC, and GCC modules.
result CSTN outperforms other methods in taxi origin-destination demand prediction.
Predicts fine-grained OD matrices for ridesharing platforms to optimize supply-demand balance.
problem Accurately predicting spatial-temporal OD demands for ridesharing platforms.
method OD-CED model combining unsupervised space coarsening and encoder-decoder architecture.
result Significant improvement in prediction accuracy (45% RMSE reduction, 60% WAPE reduction).
Study uses neural networks to predict travel times for public transportation.
problem Inaccurate travel time predictions due to road traffic irregularities.
method Developed two neural network models (MLP and LSTM) using OD travel time matrix.
result Both models can make near-accurate predictions, but LSTM is more susceptible to noise.
New properties for density-based dissimilarity measures in hybrid clustering are proposed and evaluated.
problem Choosing the right dissimilarity measure for hybrid clustering.
method Six data-independent properties for density-based dissimilarity measures are proposed and evaluated.
result A new dissimilarity measure based on Kullback-Leibler information is introduced and shown to satisfy all proposed properties.
A new model predicts dynamic O-D matrices using graph neural networks and Kalman filters.
problem Predicting dynamic O-D demand matrices from traffic flow data.
method Combines graph neural networks and Kalman filters to recognize spatial and temporal patterns.
result The proposed model outperforms other methods in various prediction scenarios.
Paper predicts in-situ metro passenger density using smart card data.
problem Crowd management in metro systems.
method Statistical models and EM algorithm for time-dependent OD matrix and travel time cost estimation.
result Accurate prediction of in-situ passenger density for future time points.
DMAE uses neural networks to cluster data with flexible dissimilarity functions.
problem Clustering data with complex dissimilarity functions.
method Integrates a dissimilarity mixture model into deep learning architectures.
result DMAE achieves competitive clustering accuracy compared to other methods.
Random Forest proximity measures for multi-view classification.
problem Combining multiple heterogeneous data views for classification.
method Building dissimilarity representations for each view, fusing them dynamically.
result Dynamic View Selection improves multi-view classification performance.
New method clusters stationary stochastic processes using covariance-based dissimilarity.
problem Clustering wide-sense stationary ergodic stochastic processes.
method Covariance-based dissimilarity measure with consistent algorithms for offline and online clustering.
result Asymptotically consistent algorithms for efficient clustering.
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
mp-LSH shares hash codes for multiple dissimilarities.
problem Hash codes depend on dissimilarity, limiting query-time adjustments.
method mp-LSH shares hash codes for L2, cosine, inner product, and weighted sums.
result mp-LSH supports user-adjustable weights and feature importance.
Proposes LOD to measure latent vs observed variables dissimilarity.
problem Quantitatively assessing relationships between latent and observed variables.
method Proposes latent-observed dissimilarity (LOD) and defines four generative model types.
result LOD effectively captures differences between models and reflects higher layer learning capability.
The paper proposes scalable methods for selecting prototypes from large dissimilarity datasets.
problem Selecting good prototypes from large dissimilarity datasets.
method Genetic algorithms, dissimilarity-based hashing, unsupervised and supervised criteria.
result The methods select good prototypes efficiently from large datasets.
The paper develops methods for clustering directed dissimilarity networks.
problem Clustering directed dissimilarity networks.
method Admissible methods based on axioms of value and transformation.
result Unique admissible clustering method exists when modifying the axiom of value.
DID measures similarity invariant to diffeomorphisms.
problem Measuring similarity invariance to diffeomorphisms.
method DID measures similarity as the solution to an optimization problem in a Reproducing Kernel Hilbert Space.
result DID is invariant to diffeomorphisms and can be efficiently approximated.
Paper introduces a new method for learning with distributions using dissimilarity measures.
problem Learning with probability distributions using dissimilarity measures.
method Introduces embeddings based on dissimilarity of distributions to templates, extending similarity theory to population distributions.
result Proves that dissimilarity theory holds for empirical distributions and shows better performance of Wasserstein distance embedding.
New dissimilarity measures enhance affinity propagation for complex network clustering.
problem Improving community detection in complex networks using affinity propagation.
method Leverage network latent geometry to design dissimilarity matrices.
result Affinity propagation outperforms state-of-the-art methods in community detection.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.
A method for classification using pairwise similarities and unlabeled data.
problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.
FPI methods compute barycenters of Gaussian sets for various dissimilarity measures.
problem Efficiently compute barycenters of Gaussian sets for multiple dissimilarity measures.
method Fixed-Point Iterations (FPI) for several dissimilarity measures.
result FPI provides a useful toolbox for fusion/reduction of Gaussian sets.
Paper proposes a new efficient transport-based dissimilarity measure for time series classification.
problem Classifying time series with warping distortions.
method Defining a problem statement, proposing an Optimal Transport-based dissimilarity measure.
result The proposed method can solve the time series classification problem with reduced computational cost.
New multitask algorithm separates rare from frequent protein functions.
problem Challenging automated protein function prediction with unbalanced data.
method Uses dissimilarity information to separate rare class labels, unlike similarity-based approaches.
result Multitask label propagation algorithm performs best with dissimilarity matrix.
ClustGeo uses Ward-like clustering with spatial constraints in R.
problem Hierarchical clustering with spatial/geographical constraints.
method Ward-like hierarchical clustering algorithm with two dissimilarity matrices and a mixing parameter.
result Determines optimal spatial contiguity without sacrificing variable quality.
A new method for functional data clustering using varying smoothing parameters.
problem Determining dissimilarity between subjects in functional data.
method Measuring dissimilarity based on varying curve estimates with commutation of smoothing parameters pair-by-pair.
result The method effectively clusters subjects and has practical advantages.
Paper introduces a method for supervised hierarchical clustering with Exponential Linkage.
problem Discrepancy between training and clustering objectives in supervised clustering.
method Tightly couples supervised training of dissimilarity function with hierarchical clustering, using Exponential Linkage.
result Joint training procedure consistently matches or outperforms other methods, improving dendrogram purity by up to 8 points.
The one-class classification problem is a well-known research endeavor in pattern recognition. The problem is also known under different names, such as outlier and novelty/anomaly detection. The core of the problem consists in modeling and recognizing patterns belonging only to a so-called target class. All other patte…
BERT improved for propaganda detection with imbalanced, dissimilar data.
problem BERT struggles with dissimilar imbalanced datasets in propaganda detection.
method Cost-sensitive BERT with dissimilarity measure for imbalanced, dissimilar datasets.
result Achieved second-highest score on sentence-level propaganda classification.
New RF dissimilarity measures improve multi-view learning accuracy.
problem Improving multi-view learning accuracy in HDLSS problems.
method Modified Random Forest proximity measures for HDLSS multi-view classification.
result Second method significantly more accurate than other state-of-the-art methods.
A new method for visualizing group dissimilarity using Likelihood Ratio Test.
problem Explaining group differences when null hypothesis is rejected.
method The Merging Path Plot and factorMerger package.
result Visualization of group dissimilarity using Likelihood Ratio Test.
Method reveals dissimilarity in alloys' Curie temperatures.
problem Tackles the dissimilarity between rare-earth transition metal binary alloys.
method Ensemble learning with Kernel ridge regression.
result Reveals meaningful relations between alloys' structure and Curie temperature.
Multiple instance learning (MIL) is concerned with learning from sets (bags) of objects (instances), where the individual instance labels are ambiguous. In this setting, supervised learning cannot be applied directly. Often, specialized MIL methods learn by making additional assumptions about the relationship of the ba…
We introduce in this paper a new way of optimizing the natural extension of the quantization error using in k-means clustering to dissimilarity data. The proposed method is based on hierarchical clustering analysis combined with multi-level heuristic refinement. The method is computationally efficient and achieves bett…
We consider the problem of identifying patterns in a data set that exhibit anomalous behavior, often referred to as anomaly detection. Similarity-based anomaly detection algorithms detect abnormally large amounts of similarity or dissimilarity, e.g.~as measured by nearest neighbor Euclidean distances between a test sam…
A new weighted dissimilarity measure reduces positioning errors in feature-based systems.
problem Reducing errors in feature-based positioning systems, especially in areas with high variability.
method Iterative scheme using location-dependent standard deviations as weights.
result Maximum radial positioning error reduced by 40% using the weighted dissimilarity measure.
FedProx algorithm improved for non-smooth and heterogeneous data.
problem Theoretical understanding of FedProx for non-convex federated optimization.
method Local dissimilarity invariant convergence theory through algorithmic stability.
result Convergence guarantees for non-smooth FL problems and minibatch size.
Method identifies regions of maximum dissimilarity in stochastic processes.
problem Comparing local characteristics of two random processes to find periods of maximum dissimilarity.
method Bayesian inference with integrated nested Laplace approximation for stochastic processes.
result Identifies regions of maximum dissimilarity with a certain volume.
In numerous applicative contexts, data are too rich and too complex to be represented by numerical vectors. A general approach to extend machine learning and data mining techniques to such data is to really on a dissimilarity or on a kernel that measures how different or similar two objects are. This approach has been …
Topolow embeds dissimilarity data into Euclidean space robustly against non-metricity and sparsity.
problem Embedding dissimilarity data into Euclidean space when dissimilarities are non-metric or sparse.
method Topolow uses a physics-inspired, gradient-free optimization framework to maximize likelihood under a Laplace error model.
result Topolow outperforms standard MDS methods in reconstructing sparse and non-Euclidean data.
New method assesses data clusterability using ultrametricity.
problem Determining if large datasets are efficiently clusterable.
method Proposes a novel ultrametric-based approach to evaluate clusterability via matrix product.
result Demonstrates the generation of sub-dominant ultrametric from dissimilarity space.
Improved detection of brain tumours in MRIs using latent space dissimilarities.
problem Detecting tumours in brain MRIs using unsupervised learning.
method Slice-wise semi-supervised method based on dissimilarity between latent representations of images and their reconstructions.
result Improved detection results with higher resolution images and better reconstructions.
The Joint Optimization of Fidelity and Commensurability (JOFC) manifold matching methodology embeds an omnibus dissimilarity matrix consisting of multiple dissimilarities on the same set of objects. One approach to this embedding optimizes the preservation of fidelity to each individual dissimilarity matrix together wi…
sCSC clusters data without prior assumptions, revealing natural groupings.
problem Clustering data without prior knowledge of its structure.
method sCSC performs binary splittings maximizing dissimilarity, producing a binary tree.
result Clusters emerge naturally from the binary tree, revealing data structure.
The paper introduces a clustering method for relational matrices using information value.
problem Clustering relational matrices of similarities or dissimilarities.
method Optimizing the value of information through a deterministic annealing process.
result The method automatically determines the number of clusters without prior specification.
Continuous MDS embeds sequences of dissimilarities in Euclidean space.
problem Embedding sequences of dissimilarities as n increases. method Continuous MDS reformulates MDS for sequences of dissimilarity matrices.
result Uniform convergence of interpolated embeddings.