Paper detects gradual changes in cluster structure using MC fusion.
problem Detecting gradual changes in cluster structure over time.
method MC fusion for multiple mixture numbers, examining MC transition.
result Accurately captures cluster structure during transitional periods.
New concept of mixture complexity helps detect gradual clustering changes.
problem Determining the number of clusters in mixture models with overlaps and weight biases.
method Introducing mixture complexity (MC) as a new measure of cluster size, defined from information theory.
result MC can detect gradual clustering changes, allowing earlier detection and finer distinction.
New RL method tackles dynamic, heterogeneous data.
problem Temporal non-stationarity and subject heterogeneity in reinforcement learning.
method Alternates between change point detection and cluster identification.
result Improves policy learning by detecting similar dynamics over time and across individuals.
Dual regularized graph Laplacian improves spectral clustering for community detection.
problem Detecting clusters in networks with improved spectral clustering methods.
method Proposes dual regularized graph Laplacian for three spectral clustering approaches.
result Theoretical analysis shows DRSC and DRSLIM yield stable consistent community detection.
A new method for real-time anomaly detection in flight data.
problem Challenges in clustering dynamically growing flight data for anomaly detection.
method Incremental Gaussian Mixture Model (GMM) using EM algorithm.
result Significantly reduced processing time and memory usage compared to offline methods.
Geometric QHD tests improve hub detection in correlated data.
problem Detecting hubs in correlated data with evolving correlations.
method Geometric QHD tests combining QCD and QHD, clustering.
result Improved hub detection in correlated data.
Most classification methods are based on the assumption that data conforms to a stationary distribution. The machine learning domain currently suffers from a lack of classification techniques that are able to detect the occurrence of a change in the underlying data distribution. Ignoring possible changes in the underly…
Detects model changes in data streams using Ddim.
problem Early detection of model changes in data streams.
method Continuous model selection based on Ddim.
result Early warning signals of model changes.
DynMSA detects market clusters for better portfolio allocation.
problem Identifying stable market clusters for effective portfolio management.
method Combining Random Matrix Theory with modularity optimization and spectral clustering.
result DynMSA outperforms baseline models in intra- and inter-cluster correlation differences.
A new method detects concept drift in streaming data using k-means space partitioning.
problem Detecting distribution changes in streaming data.
method Equal intensity k-means space partitioning (EI-kMeans) and heuristic sensitivity improvement.
result EI-kMeans improves drift detection accuracy and sensitivity.
A new method detects change points in time series with conceptors.
problem Detecting change points in time series with nonlinear temporal dependence.
method Use of conceptor matrix to learn baseline dynamics and identify change points.
result The method provides a consistent estimate of the true change point.
New method clusters financial time series into volatility regimes.
problem Finding the number of volatility regimes in nonstationary financial time series.
method Change point detection and clustering of segment distributions.
result Optimized trading strategy based on learned volatility regimes.
Sketch-based approach detects community events in evolving networks.
problem Community detection in time-varying networks.
method Maintains a small sketch graph to capture essential community structure.
result Efficiently identifies six key community events during network evolution.
New unsupervised methods for anomaly detection and clustering in structured and streaming data.
problem Anomaly detection and clustering in structured and streaming data.
method Preference Isolation Forest (PIF), Sliding-PIF, MultiLink, Online-iForest, MaxLogit.
result Methods outperform existing techniques on synthetic and real datasets.
New method detects and clusters market regimes in multidimensional data.
problem Detecting and clustering market regimes in complex data structures.
method Non-parametric online market regime detection and clustering using path-wise two-sample tests and maximum mean discrepancy.
result Successfully detected and clustered market regimes in various data structures.
We address the new problem of estimating a piece-wise constant signal with the purpose of detecting its change points and the levels of clusters. Our approach is to model it as a nonparametric penalized least square model selection on a family of models indexed over the collection of partitions of the design points and…
In this paper, we consider unsupervised partitioning problems, such as clustering, image segmentation, video segmentation and other change-point detection problems. We focus on partitioning problems based explicitly or implicitly on the minimization of Euclidean distortions, which include mean-based change-point detect…
Unified approach for non-stationary and clustered bandits.
problem Solving non-stationary and clustered bandits with overlapping solutions.
method Test of homogeneity for seamless integration of non-stationary and clustered bandits.
result Unified solution framework for change detection and cluster identification.
The article detects market regimes from covariance matrices using VLSTAR and clustering models.
problem Market regime switching is hard to detect due to time-varying correlation coefficients.
method The article applies VLSTAR and unsupervised hierarchical clustering on monthly realized covariance matrices.
result VLSTAR outperforms clustering in detecting market regimes.
We consider spectral clustering algorithms for community detection under a general bipartite stochastic block model (SBM). A modern spectral clustering algorithm consists of three steps: (1) regularization of an appropriate adjacency or Laplacian matrix (2) a form of spectral truncation and (3) a k-means type algorithm…
The vision of automated driving is to increase both road safety and efficiency, while offering passengers a convenient travel experience. This requires that autonomous systems correctly estimate the current traffic scene and its likely evolution. In highway scenarios early recognition of cut-in maneuvers is essential f…
Often the challenge associated with tasks like fraud and spam detection is the lack of all likely patterns needed to train suitable supervised learning models. This problem accentuates when the fraudulent patterns are not only scarce, they also change over time. Change in fraudulent pattern is because fraudsters contin…
We propose a new anytime hierarchical clustering method that iteratively transforms an arbitrary initial hierarchy on the configuration of measurements along a sequence of trees we prove for a fixed data set must terminate in a chain of nested partitions that satisfies a natural homogeneity requirement. Each recursive …
Method detects batch heterogeneity in genomic data.
problem Batch effects confound genomic diagnostics.
method Bayesian model evidence clustering.
result Detects batch effects without known labels.
A novel approach ODAR detects outliers for clustering.
problem Outliers interfere with clustering algorithms, leading to unreliable results.
method Feature transformation to separate outliers and normal objects into distinct clusters.
result ODAR improves clustering accuracy on 7 out of 10 datasets.
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
We investigate the tendency for financial instruments to form clusters when there are multiple factors influencing the correlation structure. Specifically, we consider a stock portfolio which contains companies from different industrial sectors, located in several different countries. Both sector membership and geograp…
The paper monitors stock market relationships using network analysis and statistical control charts.
problem Detecting abnormal changes in the financial market network structure.
method Network construction using distance methods, hierarchical clustering, and Shewhart control charts.
result Abnormal changes in financial market relationships can be detected using statistical process control.
In this paper we do the first large scale analysis of writing style development among Danish high school students. More than 10K students with more than 100K essays are analyzed. Writing style itself is often studied in the natural language processing community, but usually with the goal of verifying authorship, assess…
Improved spectral clustering for community detection in networks.
problem Community detection in networks.
method Improved spectral clustering (ISC) based on k-means clustering on weighted eigenvectors of a regularized Laplacian matrix.
result ISC yields stable consistent community detection under mild conditions and outperforms classical methods.
Develops a method to detect changes in linear systems with temporal correlations.
problem Detect abrupt changes in time series data with temporal correlations.
method Data-dependent threshold for online change point detection in linear dynamical systems.
result Achieves a pre-specified upper bound on the probability of false alarms and provides a finite-sample-based bound for detection probability.
SI-CLAD improves clustering-based anomaly detection by controlling false positives.
problem Lack of reliability in clustering-based anomaly detection.
method SI-CLAD (Statistical Inference for CLustering-based Anomaly Detection) using Selective Inference framework.
result SI-CLAD rigorously controls false detection probability below a specified significance level.
More and more neural network approaches have achieved considerable improvement upon submodules of speaker diarization system, including speaker change detection and segment-wise speaker embedding extraction. Still, in the clustering stage, traditional algorithms like probabilistic linear discriminant analysis (PLDA) ar…
We present a scheme for online, unsupervised state discovery and detection from streaming, multi-featured, asynchronous data in high-frequency financial markets. Online feature correlations are computed using an unbiased, lossless Fourier estimator. A high-speed maximum likelihood clustering algorithm is then used to f…
New algorithm detects changes in high-dimensional data with mean and variance.
problem Challenges in detecting changes in high-dimensional data with mean and variance.
method Complete graph-based approach to detect changes of mean and variance from low to high-dimensional online data.
result The proposed method outperforms existing methods in terms of detection power.
Most current clustering based anomaly detection methods use scoring schema and thresholds to classify anomalies. These methods are often tailored to target specific data sets with "known" number of clusters. The paper provides a streaming clustering and anomaly detection algorithm that does not require strict arbitrary…
Robust quickest change detection method for unknown score functions.
problem Detecting changes in data streams with unknown pre- and post-change distributions.
method Selects least-favorable distributions and robustifies score-based detection algorithm.
result Demonstrates improved performance in simulations.
New algorithm detects changes in Markov kernels with unknown post-change kernel.
problem Detecting changes in Markov kernels with unknown post-change kernel.
method Developed a new change detection algorithm assuming uniform ergodicity.
result Derived upper and lower bounds on mean delay and time between false alarms.
Bayesian methods detect clusters in noisy data more reliably.
problem Noisy data distorts traditional clustering methods, leading to unreliable results.
method Bayesian community detection using Minimum Description Length principle.
result Bayesian methods identify more robust clusters in noisy data.
New methods interpret clustering outcomes without altering data structure.
problem Post-processing methods destroy data integrity and obscure interpretations.
method Algorithm-agnostic interpretation methods using permutation feature importance, individual conditional expectation, and partial dependence.
result Preserves original feature structure and explains clustering outcomes.
Reduces change detection to estimation using confidence sequences.
problem Detecting changes in data streams with minimal delay and false alarms.
method Reduction from sequential change detection to sequential estimation using confidence sequences.
result Change detection scheme with minimal structural assumptions and strong guarantees.
Often the challenge associated with tasks like fraud and spam detection[1] is the lack of all likely patterns needed to train suitable supervised learning models. In order to overcome this limitation, such tasks are attempted as outlier or anomaly detection tasks. We also hypothesize that out- liers have behavioral pat…
NN-CUSUM detects changes in high-dimensional data using neural networks.
problem Detecting abrupt changes in high-dimensional data.
method Neural network-based CUSUM for online change-point detection.
result NN-CUSUM performs well in detecting changes in high-dimensional data.
CHAODA detects anomalies in high-dimensional data.
problem Anomaly detection in high-dimensional spaces.
method Hierarchical clustering, manifold mapping, transfer learning.
result CHAODA outperforms other algorithms on 16 out of 18 datasets.
New method detects and locates changes in spatio-temporal point processes.
problem Detecting and localizing changes in spatio-temporal data.
method Score-based, likelihood-free approach estimating change time and region.
result The method provides theoretical guarantees on detection and localization accuracy.
New algorithms detect outliers in high-dimensional data with arbitrary shapes.
problem Challenges of high dimensionality and varying cluster shapes in traditional outlier detection methods.
method Cluster Catch Digraphs (CCDs) and their variants (U-MCCD, UN-MCCD, SU-MCCD, SUN-MCCD).
result U-MCCD efficiently identifies outliers with high true negative rates, and SU-MCCD improves handling of non-uniform clusters.
Spectral clustering with edge counting detects communities in sparse models.
problem Detecting communities in sparse latent space models.
method Spectral clustering followed by edge counting.
result Algorithm achieves consistency and optimality for a broad class of models.
TS-K-means improves financial data clustering with dynamic time warping.
problem Inadequate handling of temporal dependencies in financial time series data.
method Integrates Dynamic Time Warping into Time Series K-means for financial data.
result TS-K-means outperforms traditional K-means in financial data analysis.