Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

20405979 · Jun 202019922001200920182026
48 results for dissimilarity matrices

A technique for clustering categorical data using ensembled dissimilarity matrices.

problem Clustering categorical data efficiently and accurately.
method Generate many dissimilarity matrices, average them, and extend to high dimensions using alignment techniques.
result Our method provides better clustering results, especially for genome sequences.

The paper introduces a clustering method for relational matrices using information value.

problem Clustering relational matrices of similarities or dissimilarities.
method Optimizing the value of information through a deterministic annealing process.
result The method automatically determines the number of clusters without prior specification.

New geometric structures defined on SPD matrices for better understanding.

problem Understanding SPD matrices and their geometric properties.
method Introducing Finslerian and dual information-geometric structures on James' bicone domain.
result Geodesics correspond to straight lines in coordinate systems, and new dissimilarities generalize existing ones.

ClustGeo uses Ward-like clustering with spatial constraints in R.

problem Hierarchical clustering with spatial/geographical constraints.
method Ward-like hierarchical clustering algorithm with two dissimilarity matrices and a mixing parameter.
result Determines optimal spatial contiguity without sacrificing variable quality.

New dissimilarity measures enhance affinity propagation for complex network clustering.

problem Improving community detection in complex networks using affinity propagation.
method Leverage network latent geometry to design dissimilarity matrices.
result Affinity propagation outperforms state-of-the-art methods in community detection.

A new DR method for HSI classification improves accuracy with limited samples.

problem Challenges in DR for HSI classification with limited training samples.
method Graph-based spatial and spectral regularized local scaling cut (SSRLSC).
result Improved classification accuracy compared to spectral-only methods.

CLARITY compares dissimilar datasets, identifying structural and relationship inconsistencies.

problem Integrating qualitatively different datasets from various disciplines.
method Non-parametric approach decomposing similarities into structural and relationship components.
result Identifies and interprets inconsistencies between datasets.

Proposes methods to find alternative blockmodels in networks.

problem Discover secondary blockmodel representations of networks that are dissimilar to a given blockmodel.
method Incorporates non-negative matrix factorisation (NMF) with inclusion of cannot-link constraints and dissimilarity between image matrices.
result Validated the effectiveness of the proposed methods in discovering alternative blockmodels.

A new metric learning scheme for structured data combining graph and feature-space information.

problem Learning a metric from structured data while respecting metric constraints.
method Training metric-constrained linear combinations of dissimilarity matrices, applying graph-based optimization under constraints.
result Our approach can reduce computational complexity by one order of magnitude for some cases.

Extends multidimensional scaling to analyze three-way asymmetric proximities.

problem Analyzing asymmetric and three-way proximities in a Euclidean space.
method Unified h-plot methodology for three-way asymmetric proximities, including symmetric and conditional frameworks.
result Identification of archetypal profiles and clustering structures.

New properties for density-based dissimilarity measures in hybrid clustering are proposed and evaluated.

problem Choosing the right dissimilarity measure for hybrid clustering.
method Six data-independent properties for density-based dissimilarity measures are proposed and evaluated.
result A new dissimilarity measure based on Kullback-Leibler information is introduced and shown to satisfy all proposed properties.

New method clusters stationary stochastic processes using covariance-based dissimilarity.

problem Clustering wide-sense stationary ergodic stochastic processes.
method Covariance-based dissimilarity measure with consistent algorithms for offline and online clustering.
result Asymptotically consistent algorithms for efficient clustering.

New RDPC dissimilarity measure improves time series clustering.

problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.

We have measured the dissimilarities among several printed characters of a single page in the Gutenberg 42-line bible and we prove statistically the existence of several different matrices from which the metal types where constructed. This is in contrast with the prevailing theory, which states that only one matrix per…

2010-01-31abs ↗pdf ↗

Proposes LOD to measure latent vs observed variables dissimilarity.

problem Quantitatively assessing relationships between latent and observed variables.
method Proposes latent-observed dissimilarity (LOD) and defines four generative model types.
result LOD effectively captures differences between models and reflects higher layer learning capability.

The paper proposes scalable methods for selecting prototypes from large dissimilarity datasets.

problem Selecting good prototypes from large dissimilarity datasets.
method Genetic algorithms, dissimilarity-based hashing, unsupervised and supervised criteria.
result The methods select good prototypes efficiently from large datasets.

Paper introduces a new method for learning with distributions using dissimilarity measures.

problem Learning with probability distributions using dissimilarity measures.
method Introduces embeddings based on dissimilarity of distributions to templates, extending similarity theory to population distributions.
result Proves that dissimilarity theory holds for empirical distributions and shows better performance of Wasserstein distance embedding.

Improves Gower's similarity for mixed-type variables with automatic weighting.

problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.

A method for classification using pairwise similarities and unlabeled data.

problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.

FPI methods compute barycenters of Gaussian sets for various dissimilarity measures.

problem Efficiently compute barycenters of Gaussian sets for multiple dissimilarity measures.
method Fixed-Point Iterations (FPI) for several dissimilarity measures.
result FPI provides a useful toolbox for fusion/reduction of Gaussian sets.

Paper proposes a new efficient transport-based dissimilarity measure for time series classification.

problem Classifying time series with warping distortions.
method Defining a problem statement, proposing an Optimal Transport-based dissimilarity measure.
result The proposed method can solve the time series classification problem with reduced computational cost.

New multitask algorithm separates rare from frequent protein functions.

problem Challenging automated protein function prediction with unbalanced data.
method Uses dissimilarity information to separate rare class labels, unlike similarity-based approaches.
result Multitask label propagation algorithm performs best with dissimilarity matrix.

A new method for functional data clustering using varying smoothing parameters.

problem Determining dissimilarity between subjects in functional data.
method Measuring dissimilarity based on varying curve estimates with commutation of smoothing parameters pair-by-pair.
result The method effectively clusters subjects and has practical advantages.

Paper introduces a method for supervised hierarchical clustering with Exponential Linkage.

problem Discrepancy between training and clustering objectives in supervised clustering.
method Tightly couples supervised training of dissimilarity function with hierarchical clustering, using Exponential Linkage.
result Joint training procedure consistently matches or outperforms other methods, improving dendrogram purity by up to 8 points.

The one-class classification problem is a well-known research endeavor in pattern recognition. The problem is also known under different names, such as outlier and novelty/anomaly detection. The core of the problem consists in modeling and recognizing patterns belonging only to a so-called target class. All other patte…

2014-07-28abs ↗pdf ↗

BERT improved for propaganda detection with imbalanced, dissimilar data.

problem BERT struggles with dissimilar imbalanced datasets in propaganda detection.
method Cost-sensitive BERT with dissimilarity measure for imbalanced, dissimilar datasets.
result Achieved second-highest score on sentence-level propaganda classification.

Method reveals dissimilarity in alloys' Curie temperatures.

problem Tackles the dissimilarity between rare-earth transition metal binary alloys.
method Ensemble learning with Kernel ridge regression.
result Reveals meaningful relations between alloys' structure and Curie temperature.

This work proposes a dissimilarity projection method for tractography data.

problem Tractography data cannot be directly represented in a vectorial space.
method Adopting dissimilarity representation with prototype selection and scalable approximation.
result Characterizes the use of dissimilarity projection on tractography data.

A new distance measure for HMMs using aggregated Wasserstein metric and state registration.

problem Computing dissimilarity between HMMs with Gaussian conditional distributions.
method Aggregated Wasserstein metric for Gaussian mixture distributions, state registration, optimal transport.
result The Aggregated Wasserstein distance is a semi-metric, invariant to state permutations, and more accurate than Kullback-Leibler divergence.

Multiple instance learning (MIL) is concerned with learning from sets (bags) of objects (instances), where the individual instance labels are ambiguous. In this setting, supervised learning cannot be applied directly. Often, specialized MIL methods learn by making additional assumptions about the relationship of the ba…

2013-09-22abs ↗pdf ↗

We introduce in this paper a new way of optimizing the natural extension of the quantization error using in k-means clustering to dissimilarity data. The proposed method is based on hierarchical clustering analysis combined with multi-level heuristic refinement. The method is computationally efficient and achieves bett…

2012-04-29abs ↗pdf ↗

A new weighted dissimilarity measure reduces positioning errors in feature-based systems.

problem Reducing errors in feature-based positioning systems, especially in areas with high variability.
method Iterative scheme using location-dependent standard deviations as weights.
result Maximum radial positioning error reduced by 40% using the weighted dissimilarity measure.

FedProx algorithm improved for non-smooth and heterogeneous data.

problem Theoretical understanding of FedProx for non-convex federated optimization.
method Local dissimilarity invariant convergence theory through algorithmic stability.
result Convergence guarantees for non-smooth FL problems and minibatch size.

Method identifies regions of maximum dissimilarity in stochastic processes.

problem Comparing local characteristics of two random processes to find periods of maximum dissimilarity.
method Bayesian inference with integrated nested Laplace approximation for stochastic processes.
result Identifies regions of maximum dissimilarity with a certain volume.

Topolow embeds dissimilarity data into Euclidean space robustly against non-metricity and sparsity.

problem Embedding dissimilarity data into Euclidean space when dissimilarities are non-metric or sparse.
method Topolow uses a physics-inspired, gradient-free optimization framework to maximize likelihood under a Laplace error model.
result Topolow outperforms standard MDS methods in reconstructing sparse and non-Euclidean data.

In numerous applicative contexts, data are too rich and too complex to be represented by numerical vectors. A general approach to extend machine learning and data mining techniques to such data is to really on a dissimilarity or on a kernel that measures how different or similar two objects are. This approach has been …

2014-07-02abs ↗pdf ↗