Behavioral malware clustering can be compromised by poisoning attacks.
problem Security of clustering algorithms for malware analysis.
method Investigated poisoning attacks on behavioral malware clustering.
result Behavioral malware clustering is vulnerable to poisoning attacks.
DETECT clusters mobility behaviors from trajectories using deep learning.
problem Clustering similar mobility behaviors in large, complex trajectory data.
method DETECT uses deep learning to cluster mobility behaviors from trajectories, transforming and summarizing them to identify similar behaviors.
result DETECT effectively clusters mobility behaviors from real-world datasets.
Dual-view mixture models cluster users with features and latent behaviors inferred from actions.
problem Clustering users based on features and latent behavioral functions inferred from indirect observations.
method Dual-view mixture models with non-parametric Dirichlet Process for automatic cluster number inference.
result Dual-view models outperform single-view models when one view lacks information.
Study uses contrastive learning to analyze market order behavior.
problem Understanding diverse market order behaviors.
method Self-supervised learning with triplet loss for order representation.
result Identified distinct behavior types using K-means clustering.
Unsupervised clustering identifies meal patterns in T2DM self-monitoring data.
problem Identifying individual-level behavioral-clinical phenotypes in T2DM self-monitoring data.
method Hierarchical clustering of blood glucose and macronutrient consumption.
result All 9 gold standard patterns were re-discovered using HC, and most clusters were rated positively by CDEs.
New method clusters travel behavior data from 1990-2017.
problem Challenges in analyzing large-scale travel data.
method Divide and Combine K-means clustering on time series data.
result Activity-travel patterns can be grouped into three clusters.
This study evaluates methods for clustering mobile game player behavior data.
problem Clustering time series data of player behavior in free-to-play games.
method Evaluation of various similarity measures and dimensionality reduction techniques.
result Identification and validation of temporal patterns of player behavior.
Paper uses DBSCAN variation to detect ship anomalies.
problem Detecting anomalous ship behavior.
method Variation of DBSCAN algorithm applied to AIS data.
result Alternative anomaly metric is more statistically informative.
Ridge regression shows different behaviors in binary classification with noisy labels.
problem Binary classification with noisy labels and anisotropic cluster distributions.
method Investigation of ridge regression behavior in overparameterized settings with label noise.
result Ridge regression exhibits qualitatively different behavior based on the scale of cluster mean vectors and covariance matrices.
The study shows mean partitions are consistent and asymptotically normal.
problem Lack of knowledge about the consistency of mean partitions in consensus clustering.
method Represented partitions as points in orbit space, used Fréchet means and stochastic programming, and analyzed continuous extensions of cluster criteria.
result Mean partitions are consistent and asymptotically normal under normal assumptions.
Paper models user behavior in online social media using HMMs.
problem Understanding user behavior in online social media platforms.
method Leveraging Hidden Markov Models (HMMs) to represent user behavior, deriving a model-based distance, and using spectral clustering.
result Clusters of users with similar behavioral trajectories identified.
Efficient clustering for recommender systems using bandit algorithms.
problem Improving user clustering in recommendation systems.
method Dynamic clustering based on multi-armed bandits over logged user activities.
result Enhanced prediction performance compared to existing methods.
Study evaluates clustering methods for Google Trends data.
problem Clustering high-dimensional, noisy time series data.
method Symbolic Aggregate Approximation (SAX), Enhanced SAX (eSAX), and Topological Data Analysis (TDA).
result TDA provides more balanced and meaningful groupings than SAX and eSAX.
ClusterLOB clusters market events to identify different trading behaviors.
problem Understanding market microstructure and participant behavior in financial markets.
method ClusterLOB uses K-means++ algorithm to cluster market events based on six time-dependent features.
result ClusterLOB identifies three distinct trading behaviors: directional, opportunistic, and market-making participants.
Paper presents a shape-based approach for better household load curve clustering and prediction.
problem Difficulty in classifying and predicting consumer energy consumption due to many clusters.
method Shape-based approach using Dynamic Time Warping (DTW) to align energy consumption patterns.
result Reduces the number of representative groups by 50% and improves prediction accuracy.
We propose a novel method to quantify the clustering behavior in a complex time series and apply it to a high-frequency data of the financial markets. We find that regardless of used data sets, all data exhibits the volatility clustering properties, whereas those which filtered the volatility clustering effect by using…
A k-means clustering-based SVM method classifies aggressive and moderate drivers.
problem Classifying drivers based on their curve-negotiating behaviors.
method k-means clustering for feature extraction, SVM for classification.
result kMC-SVM method reduces recognition time and improves classification accuracy.
A new deep learning method classifies VM behavior in cloud systems.
problem Scalability issues in monitoring and managing cloud data centers.
method Deep learning classifiers (DeepConv and DeepFFT) for VM behavior identification.
result Our method achieves better performance and is significantly faster.
Using geodesic currents, we provide a theoretical justification for some of the experimental results regarding the behavior of Whitehead's algorithm on non-minimal inputs, that were obtained by Haralick, Miasnikov and Myasnikov via pattern recognition methods. In particular we prove that the images of "random" elements…
Paper explores how unsupervised learning reduces financial crime risks.
problem Identifying high-risk financial groups from complex data.
method Combines clustering and dimensionality reduction techniques.
result KPCA outperforms other techniques in reducing financial crime risks.
DIVI clusters noisy high-dimensional data with stable feature gating.
problem Challenging clustering in high-dimensional noisy data.
method Data-informed variational clustering framework combining global feature gating and adaptive structure growth.
result DIVI performs competitively under severe feature noise and remains computationally feasible.
The paper addresses how clustering and prediction algorithms interact, leading to biased cross-validation results.
problem Subtle adverse behaviors in cross-validation due to interaction effects between clustering and prediction algorithms.
method The paper provides theoretical properties and scalable estimators to correct for these empirical effects.
result Expected out-of-cluster loss behavior decays rapidly with minor clustering errors, providing conditions for these effects.
Develops an adversarial clustering algorithm for detecting cyber attacks.
problem Dealing with active adversaries in cyber security data analytics.
method Grid-based adversarial clustering algorithm using game theoretic ideas.
result Identifies normal and attack objects, sub-clusters, overlapping areas, and outliers.
New method jointly clusters and learns representations for better performance.
problem Jointly clustering and learning representations for improved clustering performance.
method Continuous reparametrization of k-Means objective function. result Jointly clustering and learning representations leads to better performance.
Dynamic model clusters interactions over time, improving prediction.
problem Sparse, evolving interaction graphs with temporal dynamics.
method Structured, nonparametric edge-exchangeable model for dynamic clustering.
result Improved predictive performance compared to static models.
INCAD clusters and detects anomalies in streaming data without thresholds.
problem Clustering and anomaly detection for streaming data with unknown clusters and thresholds.
method Probabilistic clustering and anomaly detection in a streaming model.
result More reliable definition of normal vs abnormal behavior in streaming data.
New method discovers concepts in hidden feature layers using sparse subspace clustering.
problem Local attribution methods fail to identify coherent model behavior across samples.
method Sparse Subspace Clustering (SSCC) for concept discovery.
result Empirically validated method for various image classification tasks.
Study analyzes MMOG Glitch's auction house data over 14 months.
problem Understanding player behavior in MMOGs.
method Analysis of auction house data from MMOG Glitch, visualization of player migration.
result Template for analyzing player progression and churn in MMOGs.
Proposes a method to incorporate prior domain knowledge into hierarchical clustering.
problem Hierarchical clustering results depend on similarity measures and algorithm choices.
method Uses ultrametric distance function to encode external ontological information and adds it as a penalty term to the original pairwise distance.
result Popular linkage-based algorithms can faithfully recover the encoded structure.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
The paper extends cluster validity indices for incremental analysis.
problem Providing incremental alternatives for cluster validation.
method Extending iCVI family to include 6 incremental indices and examining their behavior under under- and over-partitioning.
result Over-partitioning is more challenging to detect than under-partitioning.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
This research evaluates and introduces new heuristics for clustering Bitcoin blockchain entities.
problem Efficiently analyzing the vast number of Bitcoin blockchain entities.
method Examined and introduced four new heuristics for clustering Bitcoin blockchain entities.
result Introduced clustering ratio to measure heuristic effectiveness.
The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random partitions of a subset of the sample to be dependent on the sample size, a feature not presented in a …
We extend Heston model with jumps to analyze volatility and implied volatility.
problem Analyzing implied volatility and volatility clustering in financial markets.
method Introducing an affine extension of the Heston model with α-stable jumps. result Examined jump clustering phenomenon and provided a jump cluster decomposition.
Cluster analysis of credit card accounts helps assess risk levels.
problem Assessing risk levels for credit accounts.
method Parametric modelling of account behavior, behavioral cluster analysis with a new dissimilarity measure.
result Interesting clusters and superior prediction of account default.
Clusters vehicle trajectories constrained by road networks.
problem Clustering trajectories of vehicles on road networks.
method Modelled trajectories and road segments as a bipartite graph, then clustered vertices.
result Demonstrated clustering on synthetic data, inferred flow dynamics and driver behavior.
Algorithm uncovers two main patterns of online content popularity: bursty and steady.
problem Understanding how online content gains popularity over time.
method Multi-faceted temporal analysis using dipm-SC algorithm.
result Two main patterns of popularity: bursty and steady temporal behaviors.
Modeling cascading behavior in complex systems using CTBNs.
problem Understanding which states trigger cascading events in complex systems.
method Continuous-time Bayesian networks (CTBNs) for modeling and identifying likely sentry states.
result Identification of likely sentry states that may lead to cascading behavior.
Introduces higher-order clustering coefficients to better understand network structures.
problem Understanding the clustering behavior of higher-order network cliques in complex networks.
method Develops higher-order clustering coefficients as a generalization of traditional clustering coefficients.
result Provides new insights into the structure of real-world networks.
The study explains YouTube commenters' behavior using rational inattention models.
problem Understanding and predicting YouTube commenters' behavior.
method Deep embedded clustering for user grouping, Bayesian revealed preferences for rationality testing, and behavioral economics constraints for attention span modeling.
result Most YouTube user groups optimize a Bayesian utility with rationally inattentive constraints.
Graph Ricci flow reveals hidden hierarchies in stock market correlations.
problem Detecting hidden structures in the complex stock market graph.
method Using graph Ricci curvature and flow techniques to analyze the NASDAQ 100 index.
result Algorithm detects hidden hierarchies, community behavior, and clustering in financial markets.
An empirical analysis of interest rates in money and capital markets is performed. We investigate a set of 34 different weekly interest rate time series during a time period of 16 years between 1982 and 1997. Our study is focused on the collective behavior of the stochastic fluctuations of these time-series which is in…
New bounds for convex clustering under graph connectivity.
problem Understanding clustering performance under different graph connectivity structures.
method Random walks and concentration inequalities for random graph models.
result Improved rates of convergence for centroid recovery.
The paper generalizes Thurston's earthquake map to cluster algebras of finite type.
problem Tackling Thurston's earthquake map in the context of cluster algebras of finite type.
method Introducing a cluster algebraic generalization of Thurston's earthquake map, defined by gluing exponential maps.
result Proves an analogue of the earthquake theorem for cluster algebras of finite type, showing the cluster earthquake map is a homeomorphism.
Interpolates mean shift and spectral clustering on graphs.
problem Data clustering algorithms.
method Fokker-Planck equations on data graphs.
result New theoretical insights on diffusion maps and mean shift dynamics.
Convex clustering solves a stable optimization problem for clustering.
problem Clustering with stable and scalable solutions.
method Solving a convex optimization problem with a single tuning parameter.
result The optimization problem has a unique global minimizer stable to inputs.
Solves financial volatility clustering using Minkowski metric in GARCH(1,1) model.
problem Financial volatility clustering and long memory process.
method Minkowski metric applied to GARCH(1,1) model.
result Equivalent to dark volatility or hidden risk fear field.