Spectral clustering identifies clusters of multivariate extremes.
problem Analyzing the dependence structure of multivariate extremes.
method Spectral clustering based on a random k-nearest neighbor graph. result Spectral clustering can consistently identify clusters of multivariate extremes under certain conditions.
Quantum annealing speeds up extreme clustering.
problem Efficiently grouping large datasets into many representative clusters.
method Distributed quantum annealing method.
result Optimal clustering assignments achieved under separability assumption.
Paper develops a novel approach to identify clusters of features in multivariate extremes.
problem Understanding the complex structure of multivariate extremes in various fields.
method Optimization-based approach to assess the dependence structure of extremes.
result Estimating clusters of features that best capture the support of extremes.
Kernel PCA helps analyze multivariate extremes and clusters them effectively.
problem Analyzing the dependence structure of multivariate extremes.
method Kernel PCA as a method for clustering and dimension reduction.
result Kernel PCA preimages effectively identify clusters in multivariate extremes.
APLC-XLNet improves XMTC by clustering labels and reducing computational time.
problem Efficiently tagging texts with many labels from a large set.
method Fine-tunes XLNet with APLC to approximate cross entropy loss.
result Achieved state-of-the-art results on XMTC benchmarks.
We assess cluster stability by trimming extreme points and tracking data range reduction.
problem Assessing stability of one-dimensional clusters.
method Probabilistic method using diameter-shrinkage ratio to track data range reduction.
result Our method achieves higher accuracy than classical tests in small or noisy samples.
New method clusters and visualizes anomalies in complex systems.
problem Identifying simultaneous extreme values in random vectors.
method Mixture model based on multivariate extreme value theory.
result Assigns posterior probabilities for anomaly types and clusters extreme observations.
Many modern clustering methods scale well to a large number of data items, N, but not to a large number of clusters, K. This paper introduces PERCH, a new non-greedy algorithm for online hierarchical clustering that scales to both massive N and K--a problem setting we term extreme clustering. Our algorithm efficiently …
Extreme multi-label classification aims to learn a classifier that annotates an instance with a relevant subset of labels from an extremely large label set. Many existing solutions embed the label matrix to a low-dimensional linear subspace, or examine the relevance of a test instance to every label via a linear scan. …
New method identifies key channels for extreme brain events.
problem Identifying channels responsible for extreme brain events like seizures.
method Extends canonical correlation to tail dependence, developing TPDM for clustering.
result Tail connectivity provides additional discriminatory power for seizure risk.
An explicit solution found for maximizing/minimizing agreement in a 2x2 table.
problem Maximizing or minimizing agreement between clusterings with given marginals.
method Formal framework for several agreement measures, explicit solution for 2x2 table.
result An explicit solution for the 2x2 case.
Intuitive clustering algorithm balances cluster size and cohesion.
problem Cluster definition and selection in data analysis.
method Nearest neighbours equilibrium condition for clustering.
result High-quality clustering solutions compared to benchmarks.
Formula found for minimum ARI between clusterings of fixed sizes.
problem Understanding the lowest possible agreement between clusterings.
method Explicit formula derivation for minimum ARI.
result A specific pair of clusterings achieving the minimum ARI is provided.
New risk models use chaotic attractors to predict extreme events.
problem Predicting Black Swan events in financial markets.
method Combining heavy-tailed priors with chaotic dynamics (Lorenz and Rossler systems).
result Models generate volatility clustering, fat tails, and extreme events.
In our physically inspired in-tree (IT) based clustering algorithm and the series after it, there is only one free parameter involved in computing the potential value of each point. In this work, based on the Delaunay Triangulation or its dual Voronoi tessellation, we propose a nonparametric process to compute potentia…
This paper tackles multi-modal label disentanglement in partition-based XMC.
problem Existing partition-based XMC methods create mutually exclusive clusters, which is sub-optimal for multi-modal labels.
method Formulates label assignment as an optimization problem to maximize precision rates, creating flexible and overlapped label clusters.
result Successfully disentangles multi-modal labels, leading to state-of-the-art results on XMC benchmarks.
The thesis evaluates and compares extreme mixture models in finance and insurance.
problem Estimating tail risk measures in finance and insurance.
method Extreme mixture models and methods, including kernel density estimation and GARCH preprocessing.
result Kernel density estimation-based models do not outperform others in tail risk estimation.
DCFSC uses a simple auto-encoder for subspace clustering without parameters.
problem Subspace clustering for large-scale high-dimensional datasets.
method Closed-form shallow auto-encoder for data-driven self-expressive layer, no parameters or optimization.
result Significant memory benefits over existing methods on large datasets.
In this paper we consider the problem of clustering collections of very short texts using subspace clustering. This problem arises in many applications such as product categorisation, fraud detection, and sentiment analysis. The main challenge lies in the fact that the vectorial representation of short texts is both hi…
The goal of data clustering is to partition data points into groups to minimize a given objective function. While most existing clustering algorithms treat each data point as vector, in many applications each datum is not a vector but a point pattern or a set of points. Moreover, many existing clustering methods requir…
The study identifies extremal dependence in financial markets using a bootstrap-based testing procedure.
problem Accurately identifying extremal dependence in multivariate heavy-tailed financial data.
method Bootstrap-based testing procedure applied to U.S. and Chinese stock returns.
result The U.S. exhibits more isolated clustering of dependent assets compared to China.
Proposes MSD-Kmeans for efficient outlier detection.
problem Detecting unusual records in noisy data.
method Combines MSD and K-means for more accurate outlier detection.
result MSD-Kmeans achieves highest precision, accuracy, and F-measure.
Bayesian method models normal and anomalous behaviors using extreme value theory.
problem Challenges in setting optimal thresholds for anomaly detection.
method Probabilistic framework using Dirichlet Process Mixture Model and extreme value theory.
result Explicit modeling of normal and anomalous behaviors leads to robust anomaly detection.
CEDA analyzes large categorical datasets using tree geometry and binary codes.
problem Analyzing large categorical datasets with extreme-K samples. method CEDA uses tree geometry and binary codes to analyze categorical data.
result CEDA discovers patterns and evaluates their reliability in large categorical datasets.
DEFRAG accelerates extreme classification by reducing feature dimensions.
problem High precision and scalability in assigning labels from a vast label space.
method Adaptive feature agglomeration to reduce feature dimensions.
result Significant reduction in training and prediction times (up to 40%) for extreme classification algorithms.
The study investigates the consistency of k-means clustering under finite expectation assumptions.
problem Consistency of k-means clustering under finite expectation assumptions. method Investigates the conditions under which k-means clustering is consistent, considering finite expectation instead of finite variance. result Inconsistency can arise due to extreme cluster imbalance, leading to some clusters having few points.
The extraction of clusters from a dataset which includes multiple clusters and a significant background component is a non-trivial task of practical importance. In image analysis this manifests for example in anomaly detection and target detection. The traditional spectral clustering algorithm, which relies on the lead…
Proposes a method to measure similarity between anomaly scores from different methods.
problem Difficulty in directly comparing anomaly detection methods.
method A measure based on extremal similarity in scoring distributions using a novel upper quadrant modeling approach.
result Demonstrates the ability to detect clusters of anomaly detection algorithms and achieve an accurate ensemble algorithm.
Proposes selective inference for testing differences in means between clusters.
problem Inflated type I error rate when testing differences in means between clusters.
method Selective inference approach to control selective type I error rate.
result Controls selective type I error rate by accounting for data-driven cluster definition.
A method for clustering small datasets in high dimensions using random projections.
problem Challenges in clustering small datasets in high-dimensional spaces.
method Random projection followed by binary clustering in one-dimensional space.
result Statistically significant clustering structures can be found with as few as 100-200 points.
A new algorithm removes unexpected correlations in biased data for better clustering.
problem Clustering with selection bias in data.
method Decorrelation regularized K-Means (DCKM) algorithm.
result DCKM achieves significant performance gains on real-world datasets.
Develops a new framework to measure network connectedness across and within markets.
problem Lack of flexible methods to measure network connectedness and its evolution.
method Allows network nodes to be connected in clusters, with shocks orthogonal across clusters and correlated within clusters.
result Demonstrates the effectiveness of the new framework in a detailed empirical analysis of equity markets.
This paper focuses on scalability and robustness of spectral clustering for extremely large-scale datasets with limited resources. Two novel algorithms are proposed, namely, ultra-scalable spectral clustering (U-SPEC) and ultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative selection strategy…
Paper optimizes clustering for multi-layer networks and discrete mixtures.
problem Optimizing clustering in multi-layer networks and discrete mixtures.
method Two-stage method: tensor-based initialization and likelihood-based refinement.
result Achieves minimax optimal error rate for multi-layer networks and discrete mixtures.
Model clusters authors and topics in short texts like social media posts.
problem Analysis of short texts is difficult due to brevity and lack of context.
method Expands Latent Dirichlet Allocation to model word dependencies and cluster users.
result Improves topic and user clustering, outperforming traditional methods.
We investigate task clustering for deep-learning based multi-task and few-shot learning in a many-task setting. We propose a new method to measure task similarities with cross-task transfer performance matrix for the deep learning scenario. Although this matrix provides us critical information regarding similarity betw…
We propose two spectral algorithms for partitioning nodes in directed graphs respectively with a cyclic and an acyclic pattern of connection between groups of nodes. Our methods are based on the computation of extremal eigenvalues of the transition matrix associated to the directed graph. The two algorithms outperform …
Functional data analysis is a statistical framework where data are assumed to follow some functional form. This method of analysis is commonly applied to time series data, where time, measured continuously or in discrete intervals, serves as the location for a function's value. Gaussian processes are a generalization o…
In this paper, we consider clustering data that is assumed to come from one of finitely many pointed convex polyhedral cones. This model is referred to as the Union of Polyhedral Cones (UOPC) model. Similar to the Union of Subspaces (UOS) model where each data from each subspace is generated from a (unknown) basis, in …
Transformers cluster meaningless words around leaders for sentiment analysis.
problem Capturing context in sentiment analysis using transformers.
method Characterized transformers with hardmax self-attention and normalization, showing asymptotic convergence to clustered equilibrium.
result Transformers can effectively capture context by clustering meaningless words around leader words.
Proposes ICC method for dynamic portfolio optimization.
problem Non-stationarity in market conditions makes traditional portfolio optimization ineffective.
method Inverse Covariance Clustering (ICC) to identify market states and integrate into dynamic optimization.
result ICC-PO generates portfolios with higher Sharpe Ratios and greater robustness.
The paper introduces a new model to improve exotic option pricing.
problem Challenges in pricing exotic options and structured products due to market phenomena.
method Introduces a Diffusion-Conditional Probability Model (DDPM) with a composite loss function and P-Q dynamic game framework.
result The DDPM outperforms traditional models in dynamic games for European and Asian options, but underestimates tail risks.
Range penalization enhances statistical accuracy and resource efficiency in federated learning.
problem Statistical accuracy and resource efficiency in federated learning.
method Range regularization and polar clustering.
result Enhanced statistical accuracy and reduced iteration complexity.
This paper investigates the model degrees of freedom in k-means clustering. An extension of Stein's lemma provides an expression for the effective degrees of freedom in the k-means model. Approximating the degrees of freedom in practice requires simplifications of this expression, however empirical studies evince the a…
Unified approach for clustering financial multiplex networks.
problem Lack of methods to capture interconnections between assets over time.
method Tensor-based unified local and global clustering coefficients for multiplex networks.
result Unified clustering coefficients effectively describe dependencies between assets over time.
Study of 2D Ising model reveals patterns in financial markets.
problem Understanding stylized facts in financial markets using statistical physics.
method 2D Ising model with spin interactions; analysis of spin clusters, persistence, and dynamics.
result Microscopic mechanisms explain stylized facts like sharp peaks in returns and heavy-tailed distributions.
A new multi-view clustering method that is fast, scalable, and easy to use.
problem High computational complexity, one-stage fusion, and dataset-specific hyperparameter tuning in multi-view clustering.
method Random view groups, hybrid early-late fusion, diversified base clusterings, and unified bipartite graph.
result Almost linear time and space complexity, no dataset-specific tuning required.
We build a simple model of leveraged asset purchases with margin calls. Investment funds use what is perhaps the most basic financial strategy, called "value investing", i.e. systematically attempting to buy underpriced assets. When funds do not borrow, the price fluctuations of the asset are normally distributed and u…