A novel method relaxes binary constraints to non-negative spheres for multi-matching and clustering.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.
We consider the problem of universal joint clustering and registration of images and define algorithms using multivariate information functionals. We first study registering two images using maximum mutual information and prove its asymptotic optimality. We then show the shortcomings of pairwise registration in multi-i…
Clustering is one of the most universal approaches for understanding complex data. A pivotal aspect of clustering analysis is quantitatively comparing clusterings; clustering comparison is the basis for many tasks such as clustering evaluation, consensus clustering, and tracking the temporal evolution of clusters. In p…
NOs can learn any finite collection of classes in functional data.
New criterion selects optimal number of clusters based on stability.
The problem of universal outlying sequence detection is studied, where the goal is to detect outlying sequences among sequences of samples. A sequence is considered as outlying if the observations therein are generated by a distribution different from those generating the observations in the majority of the sequenc…
Subspace clustering refers to the problem of clustering high-dimensional data into a union of low-dimensional subspaces. Current subspace clustering approaches are usually based on a two-stage framework. In the first stage, an affinity matrix is generated from data. In the second one, spectral clustering is applied on …
Consider unsupervised clustering of objects drawn from a discrete set, through the use of human intelligence available in crowdsourcing platforms. This paper defines and studies the problem of universal clustering using responses of crowd workers, without knowledge of worker reliability or task difficulty. We model sto…
A clustering algorithm for natural hierarchical clusters with near-linear time complexity.
Subspace clustering (SC) refers to the problem of clustering high-dimensional data into a union of low-dimensional subspaces. Based on spectral clustering, state-of-the-art approaches solve SC problem within a two-stage framework. In the first stage, data representation techniques are applied to draw an affinity matrix…
Clusters of highly correlated stocks are identified for better asset selection.
TACE unifies scalar and tensorial modeling in Cartesian space for accurate, stable, and efficient atomistic predictions.
Method detects lead-lag relationships in multivariate time series.
The paper uses clustering and integer programming to optimize stock selection for investment funds.
Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often makes the algorithm selection and hyperparameter evaluation a tough guess. In t…
NeuralFLoC unifies registration and clustering of functional data, overcoming phase variation challenges.
We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…
Financial price changes obey two universal properties: they follow a power law and they tend to be clustered in time. The second regularity, known as volatility clustering, entails some predictability in the price changes: while their sign is uncorrelated in time, their amplitude (or volatility) is long-range correlate…
New method clusters directed graphs using Koopman operators.
A novel clustering method uses torque balance to group objects.
Benchmark study evaluates 8 clustering methods on 99 UCR time series datasets.
The present work proposes hybridization of Expectation-Maximization (EM) and K-Means techniques as an attempt to speed-up the clustering process. Though both K-Means and EM techniques look into different areas, K-means can be viewed as an approximate way to obtain maximum likelihood estimates for the means. Along with …
Clusters of crypto assets by path signature improve diversification and reduce fees.
Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.
In this paper, we present a novel way to summarize the structure of large graphs, based on non-parametric estimation of edge density in directed multigraphs. Following coclustering approach, we use a clustering of the vertices, with a piecewise constant estimation of the density of the edges across the clusters, and ad…
With pressure to increase graduation rates and reduce time to degree in higher education, it is important to identify at-risk students early. Automated early warning systems are therefore highly desirable. In this paper, we use unsupervised clustering techniques to predict the graduation status of declared majors in fi…
We propose an efficient Context-Aware clustering of Bandits (CAB) algorithm, which can capture collaborative effects. CAB can be easily deployed in a real-world recommendation system, where multi-armed bandits have been shown to perform well in particular with respect to the cold-start problem. CAB utilizes a context-a…
Model predicts higher education dropout risk with interpretable parameters.
Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between them, or a mix of both. Existing methods for relational clustering have strong and …
MPE framework proves universal approximation for quantum data distribution.
A new CVI called DSI evaluates clustering results without true labels.
We consider the task of estimating a Gaussian graphical model in the high-dimensional setting. The graphical lasso, which involves maximizing the Gaussian log likelihood subject to an l1 penalty, is a well-studied approach for this task. We begin by introducing a surprising connection between the graphical lasso and hi…
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.
Extracting significant places or places of interest (POIs) using individuals' spatio-temporal data is of fundamental importance for human mobility analysis. Classical clustering methods have been used in prior work for detecting POIs, but without considering temporal constraints. Usually, the involved parameters for cl…
Market sectors play a key role in the efficient flow of capital through the modern Global economy. We analyze existing sectorization heuristics, and observe that the most popular - the GICS (which informs the S&P 500), and the NAICS (published by the U.S. Government) - are not entirely quantitatively driven, but rather…
Researchers improve Gaussian processes to model inconsistent preferences.
The study examines the universality of Gaussian data in high-dimensional generalized linear estimation.
Quantum cluster algebra constructed from web skein relations on surfaces.
Transformers learn to cluster Gaussian mixtures as well as the EM algorithm.
We introduce the cluster exchange groupoid associated to a non-degenerate quiver with potential, as an enhancement of the cluster exchange graph. In the case that arises from an (unpunctured) marked surface, where the exchange graph is modelled on the graph of triangulations of the marked surface, we show that the univ…
Study of 2D Ising model reveals patterns in financial markets.
Unified HDP and LDA models for efficient topic clustering of online course queries.
Algorithm refines matrix ratings using hierarchical graph clustering.
The Minimal Learning Machine (MLM) is a nonlinear supervised approach based on learning a linear mapping between distance matrices computed in the input and output data spaces, where distances are calculated using a subset of points called reference points. Its simple formulation has attracted several recent works on e…
The large-scale structure of the universe is comprised of virialized blob-like clusters, linear filaments, sheet-like walls and huge near empty three-dimensional voids. Characterizing the large scale universe is essential to our understanding of the formation and evolution of galaxies. The density range of clusters, wa…
Modeling videos and image-sets as linear subspaces has proven beneficial for many visual recognition tasks. However, it also incurs challenges arising from the fact that linear subspaces do not obey Euclidean geometry, but lie on a special type of Riemannian manifolds known as Grassmannian. To leverage the techniques d…