Unified framework for efficient Frank-Wolfe optimization of Dominant Set Clustering.
problem Optimizing Dominant Set Clustering with various Frank-Wolfe algorithms.
method Unified framework for pairwise, standard, and away-steps Frank-Wolfe algorithms, with explicit convergence rates.
result Explicit convergence rates for Frank-Wolfe methods in Dominant Set Clustering.
This paper proposes a new clustering method based on Stochastic Dominance for asset allocation.
problem Traditional clustering methods fail to capture risk dominance relationships among assets.
method Integrates Stochastic Dominance theory with machine learning algorithms to construct a Stochastic Dominance Coefficient Matrix and modify clustering algorithms.
result The proposed method effectively facilitates customized asset allocation for investors.
The increasing needs of clustering massive datasets and the high cost of running clustering algorithms poses difficult problems for users. In this context it is important to determine if a data set is clusterable, that is, it may be partitioned efficiently into well-differentiated groups containing similar objects. We …
We propose clustering algorithms based on a recently developed geometric digraph family called cluster catch digraphs (CCDs). These digraphs are used to devise clustering methods that are hybrids of density-based and graph-based clustering methods. CCDs are appealing digraphs for clustering, since they estimate the num…
We propose a combination of cluster analysis and stochastic process analysis to characterize high-dimensional complex dynamical systems by few dominating variables. As an example, stock market data are analyzed for which the dynamical stability as well as transitions between different stable states are found. This comb…
Paper uses ANFIS to predict cryptocurrency prices.
problem Predicting cryptocurrency prices for seven days.
method Adaptive Network Based Fuzzy Inference System (ANFIS) with hybrid and backpropagation algorithms.
result The method can predict cryptocurrency prices in a short time.
A new deep clustering model learns both clustering and embedding simultaneously.
problem Optimizing autoencoders for clustering and embedding separately.
method Integrating clustering module into a deep autoencoder.
result Joint learning of clustering and embedding improves performance.
Identifies key industrial sectors in S&P 500 states.
problem Understanding changing market conditions in financial markets.
method Clustering algorithm, XAI relevance scores, Bayesian change point analysis.
result Dominant sectors (energy and IT) determine market states.
New method for selecting clusters in residential electricity data.
problem Selecting useful clusters in electricity consumption data.
method Formalizing expert knowledge as external validation measures.
result Successfully reconstructed customer archetypes.
Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a collection of about 30,000 tweets extracted from Twitter just before the World Cup st…
We extend the theoretical analysis of a recently proposed single subspace learning algorithm, called Dual Principal Component Pursuit (DPCP), to the case where the data are drawn from of a union of hyperplanes. To gain insight into the properties of the ℓ1 non-convex problem associated with DPCP, we develop a geo…
Connected domination numbers found for plane triangulations up to 13 vertices.
problem Finding connected domination numbers for plane triangulations.
method Analyzing triangulations of up to 13 vertices and proving the difference between connected and regular domination numbers can be arbitrarily large.
result Connected domination numbers for triangulations up to 13 vertices and upper bound for larger triangulations.
Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the…
The paper explores fair clustering, a niche area in machine learning.
problem Fairness in clustering remains underexplored despite its importance.
method Assesses existing work and proposes new directions for fair clustering.
result Widening normative principles and knowledge of downstream processes can enhance fair clustering research.
FOSC-X: An extended framework for extracting multiple optimal flat clusterings from hierarchical cluster trees
problem Extracting multiple optimal flat clusterings from hierarchical cluster trees
method Dynamic programming with lower and upper feasibility bounds
result Guaranteed optimal rankings of top-M solutions with linear-time complexity
Proposes a new clustering method based on expectiles for non-spherical clusters.
problem Inability of K-means to handle non-spherical clusters. method Uses expectiles to define cluster centers and searches for clusters via a greedy algorithm.
result Outperforms K-means and spectral clustering on asymmetric shaped clusters. Paper proposes a new method to improve clustering ensemble performance.
problem Improving clustering ensemble performance by refining co-association matrix.
method Low-rank tensor approximation to derive coherent-link matrix and refine co-association matrix.
result The proposed method achieves breakthrough in clustering performance compared to state-of-the-art methods.
MAS scores cluster size consistency from points, robust to label changes.
problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.
This paper presents a neural network-based end-to-end clustering framework. We design a novel strategy to utilize the contrastive criteria for pushing data-forming clusters directly from raw data, in addition to learning a feature embedding suitable for such clustering. The network is trained with weak labels, specific…
Baseline injury categorization is important to traumatic brain injury (TBI) research and treatment. Current categorization is dominated by symptom-based scores that insufficiently capture injury heterogeneity. In this work, we apply unsupervised clustering to identify novel TBI phenotypes. Our approach uses a generaliz…
Complex network analysis reveals dominant stocks in financial stock returns correlations.
problem Inferring financial stock returns correlations from complex network analysis.
method Simulated geometric Brownian motion for stocks, complex network analysis, eigenvector centrality, clustering.
result Returns correlation matrix is dominated by stocks with high eigenvector centrality and clustering.
The current trend of pushing CNNs deeper with convolutions has created a pressing demand to achieve higher compression gains on CNNs where convolutions dominate the computation and parameter amount (e.g., GoogLeNet, ResNet and Wide ResNet). Further, the high energy consumption of convolutions limits its deployment on m…
Entropy regularization improves interpretability of probabilistic clustering models.
problem Bayesian nonparametric mixture models often produce unbalanced cluster frequencies.
method Interpreting the posterior as penalized likelihood, entropy regularization reduces sparsely-populated clusters.
result The proposed entropy-regularized estimator enhances interpretability without sacrificing computational convenience.
New method separates market motion from stock correlations.
problem Understanding the dynamics of stock correlations relative to market motion.
method Cluster reduced-rank correlation matrices by subtracting the largest eigenvalue.
result Extracted market states are quasi-stationary over long periods.
Framework clusters noisy MTS with robust fuzzy clustering, improving accuracy over existing methods.
problem Challenges in clustering multivariate time series due to non-stationary dependencies, noise, and state boundaries.
method Spectral fuzzy clustering using Kendall's tau-based canonical coherence for frequency-specific monotonic relationships.
result Framework outperforms existing methods in clustering noisy, high-dimensional MTS.
This paper considers the problem of estimating a high-dimensional vector of parameters θ∈Rn from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss function, the James-Stein (JS) estimator is known to dominate the simple maximum-likelihood (…
Genie clusters faster and resists outliers.
problem Hierarchical clustering's sensitivity to outliers and slow computation.
method Genie uses an economic inequity measure to link clusters, balancing speed and quality.
result Genie outperforms other linkage methods in clustering quality and speed.
Proves rigidity for specific initial data sets under the dominant energy condition.
problem Rigidity of initial data sets with boundary and convex polytopes.
method Solution of boundary value problems for Dirac operators and approximations by manifolds with smooth boundary.
result Proves rigidity for compact smooth spin manifolds and convex polytopes under the dominant energy condition.
The study uses unsupervised machine learning to identify top European football teams.
problem Selecting teams for the new European football Super League.
method Used Laplacian eigenmaps clustering on performance data.
result Successfully identified four clusters of teams based on performance metrics.
New method ranks multivariate distributions in SMOOP using q-dominance.
problem Lack of reliable methods to rank multivariate distributions in SMOOP.
method Introduces center-outward q-dominance and develops empirical test procedures.
result Proves q-dominance implies FSD and establishes a sample size threshold.
New homology theory connects graph domination to subtle algebraic structures.
problem Understanding graph domination through algebraic homology.
method Interpreting überhomology as poset homology and showing its functorial properties.
result The Euler characteristic of bold homology equals the evaluation of the connected domination polynomial.
Kernel-based K-means clustering has gained popularity due to its simplicity and the power of its implicit non-linear representation of the data. A dominant concern is the memory requirement since memory scales as the square of the number of data points. We provide a new analysis of a class of approximate kernel methods…
We consider localized deformation for initial data sets of the Einstein field equations with the dominant energy condition. Deformation results with the weak inequality need to be handled delicately. We introduce a modified constraint operator to absorb the first order change of the metric in the dominant energy condit…
In this work we introduce two novel deterministic annealing based clustering algorithms to address the problem of Edge Controller Placement (ECP) in wireless edge networks. These networks lie at the core of the fifth generation (5G) wireless systems and beyond. These algorithms, ECP-LL and ECP-LB, address the dominant …
In this paper, we focus on the stochastic block model (SBM),a probabilistic tool describing interactions between nodes of a network using latent clusters. The SBM assumes that the networkhas a stationary structure, in which connections of time varying intensity are not taken into account. In other words, interactions b…
Proves positive mass theorem for spin initial data sets with arbitrary ends and dominant energy shields.
problem Proving the positive mass theorem for spin initial data sets with various ends and energy shields.
method Modification of Witten's approach involving an additional independent timelike direction in the spinor bundle.
result Positive mass theorem for spin initial data sets with arbitrary ends and dominant energy shields.
Hierarchical clustering is a class of algorithms that seeks to build a hierarchy of clusters. It has been the dominant approach to constructing embedded classification schemes since it outputs dendrograms, which capture the hierarchical relationship among members at all levels of granularity, simultaneously. Being gree…
Odd-dimensional manifolds have contact maps of non-zero degree.
problem Contact domination in odd-dimensional manifolds.
method Proving existence of maps from tight contact manifolds.
result Existence of non-zero degree maps from Liouville-fillable but not Weinstein-fillable contact manifolds.
New insights into Bartnik mass from improvability of dominant energy scalar.
problem Characterizing Bartnik mass minimizing initial data sets.
method Introducing improvability concept, proving non-improvability consequences, and analyzing pp-wave counterexamples.
result Bartnik mass minimizing initial data sets are characterized, advancing conjectures.
New framework for ranking distributions using variable fractional parameters.
problem Ordering distributions with varying steepness and local non-concavities.
method Introducing a function γ:Ro[0,1] to replace the fixed parameter in fractional SD. result Enables ranking of a broader range of distributions and incorporates dynamic greediness.
Proposes a new method to learn data representations by modeling sample relations.
problem Lack of rich latent structural information in DAEs.
method Explicitly models and leverages sample relations as supervision for representation learning.
result Significantly improves clustering performance on benchmark datasets.
Automatic cover detection -- the task of finding in a audio dataset all covers of a query track -- has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to detect automatically if an audio excerpt embeds musical content belonging to the…
Smooth dec initial data sets may not extend to smooth spacetimes.
problem Whether every dec initial data set can be extended to a smooth spacetime.
method Examined the converse of the dominant energy condition for initial data sets and spacelike hypersurfaces.
result Not all dec initial data sets can be extended to smooth spacetimes.
X-DC improves speech separation by making DNNs more interpretable.
problem Black-box nature of DNNs in speech separation tasks.
method Introduces X-DC, a DNN architecture that interprets as spectrogram template fitting followed by Wiener filtering.
result X-DC achieves comparable speech separation performance to DC but with enhanced interpretability.
New algorithm finds sparse matrices on Stiefel manifold for optimisation.
problem Finding sparse matrices on Stiefel manifold for optimisation.
method Modified Orthogonal Iteration algorithm for sparse global optimality.
result Proposed method finds globally optimal sparse Stiefel matrices.
SC-InfoNCE improves InfoNCE for feature clustering in contrastive learning.
problem Lack of theoretical understanding of InfoNCE's feature clustering mechanism.
method Introduced a transition probability matrix to model data augmentation dynamics and optimize feature similarity.
result SC-InfoNCE achieves strong performance across diverse domains, aligning feature similarity with downstream data.
Characterizes causal structure dominance for latent variables.
problem Determining dominance relations between causal structures with latent variables.
method Complete characterization for three visible variables, partial for four; uses nontrivial inequality constraints.
result Equivalence classes with nontrivial inequality constraints become ubiquitous as the number of visible variables increases.
Proves density and mass theorems for specific initial data sets.
problem Initial data sets with boundary in spacetime.
method Harmonic asymptotics and dominant energy condition.
result Spacetime positive mass theorem for initial data sets with apparent horizon boundary.