Formula found for minimum ARI between clusterings of fixed sizes.
problem Understanding the lowest possible agreement between clusterings.
method Explicit formula derivation for minimum ARI.
result A specific pair of clusterings achieving the minimum ARI is provided.
Proposes new random models for fuzzy clustering similarity measures.
problem Challenges in choosing a random model for fuzzy clustering similarity measures.
method Introduces three intuitive and explainable random models for fuzzy clusterings.
result Each random model has distinct behavior, emphasizing the importance of accurate model selection.
The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to better understand exactly what they measure, their properties and their differences. …
Meta-learning neural networks for better clustering representations.
problem Improving clustering performance with appropriate representations.
method Meta-learning method that trains neural networks for representations using VB inference with an infinite Gaussian mixture model.
result The method achieves higher clustering performance than existing methods.
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.
problem Divergent evaluations of unsupervised models due to different clustering similarity measures.
method Developed an analytical framework that unifies pair-counting and information-theoretic clustering similarity measures.
result Unified framework clarifies when and why the two regimes diverge and provides a principled basis for selecting and interpreting clustering similarity measures.
The goal of lifetime clustering is to develop an inductive model that maps subjects into K clusters according to their underlying (unobserved) lifetime distribution. We introduce a neural-network based lifetime clustering model that can find cluster assignments by directly maximizing the divergence between the empiri…
The main goal of this study is to extract a set of brain networks in multiple time-resolutions to analyze the connectivity patterns among the anatomic regions for a given cognitive task. We suggest a deep architecture which learns the natural groupings of the connectivity patterns of human brain in multiple time-resolu…
New metric improves clustering in persistent homology.
problem Improving clustering accuracy in persistent homology.
method Defined a new non-archimedean cophenetic metric.
result Cophenetic metric enhances clustering quality and inter-relations.
Paper addresses xVA models for market-implied skew and smile.
problem Capturing market-implied skew and smile in xVA calculations.
method Developed a state-dependent SDE combining Hull-White models with RAnD technique.
result Demonstrated significant effect of skew and smile on xVA calculations.
A new measure DCSI quantifies separability for density-based clustering.
problem Quantifying meaningful clusters in data sets.
method Developed a new separability measure DCSI based on separation and connectedness.
result Correctly identifies touching or overlapping classes that do not correspond to meaningful density-based clusters.
Randomizes AD models for better option pricing.
problem Inconsistent option pricing with affine models.
method Randomization of AD models with exogenous stochasticity.
result RAnD models allow for better calibration and consistent pricing.
It has been noticed that some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand index (R…
Unsupervised image segmentation aims at clustering the set of pixels of an image into spatially homogeneous regions. We introduce here a class of Bayesian nonparametric models to address this problem. These models are based on a combination of a Potts-like spatial smoothness component and a prior on partitions which is…
In mixture model-based clustering applications, it is common to fit several models from a family and report clustering results from only the `best' one. In such circumstances, selection of this best model is achieved using a model selection criterion, most often the Bayesian information criterion. Rather than throw awa…
CDL index improves clustering validation for non-convex data.
problem Selecting clustering algorithms and hyperparameters without labeled data.
method CDL uses compactness, centers, and covariances to compute a probabilistic description length bound.
result CDL outperforms conventional CVIs on synthetic and image benchmarks.
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
Study EM and GD for clustering with penalties for misspecification and high dimensions.
problem Clustering with misspecification and high-dimensional data.
method Model-based Gaussian Mixture Models, EM algorithm, GD optimization with AD, penalized likelihood.
result GD outperforms EM on high-dimensional data but both have poor cluster interpretation.
A new clustering framework optimizes customer search data for personalized travel recommendations.
problem Personalized travel recommendations based on customer search data.
method Multi-objective optimization-based clustering ensemble framework.
result Optimizes diversity in clustering ensemble search space and automatically determines the number of clusters.
Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open …
STICC clusters geographic objects considering both spatial contiguity and attributes.
problem Discovering repeated geographic patterns with spatial contiguity.
method Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method.
result STICC significantly outperforms baseline methods in adjusted rand index and macro-F1 score.
Mixtures of Unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by Multinomials. When the classification task is particularly challenging, such as when the document-term ma…
Study compares clustering methods for mixed-type data.
problem Challenges in clustering mixed-type data.
method Distance-based (k-prototypes, PDQ, convex k-means), probabilistic (KAY-means, MBNs, LCM).
result KAMILA, LCM, and k-prototypes perform best.
A new measure normalizes clustering accuracy to evaluate algorithms better.
problem Evaluation of clustering algorithms is challenging due to limitations of existing measures.
method Proposes a new, normalised clustering accuracy measure.
result The new measure identifies worst-case scenarios and is more interpretable.
FCM clustering adapts to persistence diagrams for topological data analysis.
problem Integrating topological data into machine learning workflows.
method Adapting Fuzzy c-Means to persistence diagrams.
result FCM clustering captures topological structure without additional processing.
Clustering is a central approach for unsupervised learning. After clustering is applied, the most fundamental analysis is to quantitatively compare clusterings. Such comparisons are crucial for the evaluation of clustering methods as well as other tasks such as consensus clustering. It is often argued that, in order to…
Improved zeroth-order algorithms for nonconvex optimization with reduced complexity and improved performance.
problem Designing efficient zeroth-order algorithms for nonconvex optimization with reduced function query complexities and improved convergence rates.
method Proposed new algorithms ZO-SVRG-Coord-Rand and ZO-SPIDER-Coord, developed new analyses, and addressed issues of function query complexities and stepsize generation.
result New algorithms outperform existing methods in terms of function query complexities and convergence rates.
New k-means method handles random data better than traditional techniques.
problem Limitations of traditional clustering methods in random data.
method Probabilistic metric space with random normed k-means (RNKM).
result RNKM outperforms traditional methods in complex clustering scenarios.
The study of genetic variants can help find correlating population groups to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning algorithms are increasingly being applied to identify interacting GVs to understand th…
Visual summarization of clinical data collected on patients contained within the electronic health record (EHR) may enable precise and rapid triage at the time of patient presentation to an emergency department (ED). The triage process is critical in the appropriate allocation of resources and in anticipating eventual …
The paper investigates how irrelevant features affect clustering performance.
problem The challenge of identifying relevant features in unsupervised clustering tasks.
method Investigation of clustering performance with added irrelevant features.
result Different types of irrelevant features impact clustering outcomes differently.
New method clusters matrix-valued data by latent variables.
problem Clustering matrix-valued data with hidden structure.
method Latent variable model with hierarchical clustering.
result Algorithm attains clustering consistency in high dimensions.
The paper introduces a new model to correct bias in treatment effect estimates due to sample selection.
problem Bias in treatment effect estimates due to sample selection.
method Type 2 Tobit Bayesian Additive Regression Trees (TOBART-2) with Dirichlet Process Mixture distribution and soft trees.
result Corrects bias in treatment effect estimates by accounting for nonlinearities and model uncertainty.
The paper critiques and expands on common evaluation metrics in machine learning.
problem The common evaluation metrics like Precision, Recall, F-Measure, and Rand Accuracy are biased and misleading.
method The paper introduces new measures like Informedness, Markedness, and Correlation to better reflect the quality of predictions.
result A system that performs worse in terms of Informedness can appear better using common measures like Precision and Recall.
We enhance short-rate models to control implied volatility analytically.
problem Controlling implied volatility in short-rate models.
method Randomized Affine Diffusion (RAnD) method applied to Heath-Jarrow-Morton framework.
result Randomized short-rate models improve calibration and control implied volatility shapes.
A new clustering evaluation index based on density estimation.
problem Improving internal clustering evaluation indices.
method The index is a mixture of Ambiguous and Similarity sub-indices, calculated using density estimation.
result The new index significantly outperforms other internal clustering evaluation indices.
Study on symmetric operators on non-compact manifolds, focusing on their index modulo 2.
problem Investigating elliptic operators with a specific symmetry and their index modulo 2.
method Analysis of Callias-type operators on non-compact manifolds, establishing mod 2 versions of index theorems.
result Established mod 2 versions of the Gromov-Lawson relative index theorem, Callias index theorem, and Boutet de Monvel's index theorem for Toeplitz operators.
New methods reduce communication in distributed training for variational inequalities.
problem Reducing communication in distributed training for high-dimensional models.
method Distributed methods with compressed communication for solving variational inequalities.
result Theoretical guarantees and practical algorithms for compressed communication.
New index formula connects numerical and K-theoretic indices.
problem Equivariant index for proper group actions on manifolds.
method Developed a trace on group conjugacy classes to relate numerical and K-theoretic indices. result Shows that numerical index equals K-theoretic index under certain conditions. The paper explores global index formulas for one-dimensional holomorphic foliations.
problem Global index formulas for one-dimensional holomorphic foliations.
method Microlocal point of view and short proofs for existing index formulas.
result Generalizations of existing index formulas.
Explain Arnold's proof of the Morse index theorem using Maslov index.
problem Proving the Morse index theorem in Riemannian geometry.
method Using symplectic arguments and the Maslov index.
result Self-contained exposition of Arnold's proof.
The p-index improves investment performance for NYSE stocks but not for SSE stocks.
problem Improving investment performance for stocks using the p-index.
method Comparing different p-ratio strategies and empirical efficient frontiers for SSE and NYSE stocks.
result The p-index enhances investment performance for NYSE stocks but not for SSE stocks.
Paper introduces danceability index as a new bridge index definition.
problem Defining the bridge index in various mathematical contexts.
method Proves danceability index as equivalent to bridge index, extends to virtual knots.
result Danceability index is a new equivalent definition of the bridge index.
Study Whittle index learning algorithms for restless bandits with constant stepsizes.
problem Optimizing decisions in restless multi-armed bandits with constant stepsizes.
method Developed Q-learning algorithms with constant stepsizes for index learning in restless bandits, extending to DQN and function approximations.
result The algorithms learn the Whittle index effectively.
A new index rebalancing strategy reduces large constituent weights without undesirable effects.
problem Undesirable effects of current Nasdaq-100 index rebalancing.
method A simple rebalancing strategy that avoids undesirable effects.
result Preserves the order of index weights and prevents maximum weight increase.
We study bounded pseudoconvex domains in complex Euclidean space. We define an index associated to the boundary and show this new index is equivalent to the Diederich-Fornæss index defined in 1977. This connects the Diederich-Fornæss index to boundary conditions and refines the Levi pseudoconvexity. We also prove the $…
Study on symmetric braid index of ribbon knots, deriving bounds and characterizations.
problem Understanding the symmetric braid index of ribbon knots.
method Defining symmetric braid index, using Khovanov homology, and calculating bounds.
result Existence of knots with symmetric braid index greater than braid index.
Study proves bridge and braid indices match for twist positive knots.
problem Determining when bridge and braid indices are equal for knots.
method Used knot Floer torsion order to prove for all twist positive knots.
result Bridge and braid indices coincide for all twist positive knots.