We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy c…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many modern clustering methods scale well to a large number of data items, N, but not to a large number of clusters, K. This paper introduces PERCH, a new non-greedy algorithm for online hierarchical clustering that scales to both massive N and K--a problem setting we term extreme clustering. Our algorithm efficiently …
Graph auto-encoders improve financial clustering using news and stock data.
Graph learning categorizes DeFi services into similar functionalities.
In this work we propose a simple and easily parallelizable algorithm for multiway graph partitioning. The algorithm alternates between three basic components: diffusing seed vertices over the graph, thresholding the diffused seeds, and then randomly reseeding the thresholded clusters. We demonstrate experimentally that…
New subspace prototype flag median improves clustering on noisy data.
In supervised clustering, standard techniques for learning a pairwise dissimilarity function often suffer from a discrepancy between the training and clustering objectives, leading to poor cluster quality. Rectifying this discrepancy necessitates matching the procedure for training the dissimilarity function to the clu…
Clustering analysis by nonnegative low-rank approximations has achieved remarkable progress in the past decade. However, most approximation approaches in this direction are still restricted to matrix factorization. We propose a new low-rank learning method to improve the clustering performance, which is beyond matrix f…
Improves hierarchical clustering in Euclidean space using autoencoders.
All the connections, pure toward the nilpotent structure, are found. Examples of manifolds, for which the curvature tensor is pure or hybrid, are given. For a manifold of B-type a necessary and sufficient condition for purity of the curvature tensor is proved. It is verified that the conformal change of the metric of a…
We propose an effective method to solve the event sequence clustering problems based on a novel Dirichlet mixture model of a special but significant type of point processes --- Hawkes process. In this model, each event sequence belonging to a cluster is generated via the same Hawkes process with specific parameters, an…
This paper examines how regional trade agreements affect global trade relationships.
Point source detection at low signal-to-noise is challenging for astronomical surveys, particularly in radio interferometry images where the noise is correlated. Machine learning is a promising solution, allowing the development of algorithms tailored to specific telescope arrays and science cases. We present DeepSourc…
We provide a new local class-purity theorem for Lipschitz continuous DNN classifiers. In addition, we discuss how to achieve classification margin for training samples. Finally, we describe how to compute margin p-values for test samples.
The paper proposes a method to identify high-quality financial patterns using entropy.
QNA uses quantum-inspired density operators to diagnose market dependence and structural risk.
Deep learning model creates patient representations for scalable EHR-based stratification.
Geometric observables detect financial regime shifts with high accuracy.
The study classifies normal subgroups of mapping class groups of surfaces with Cantor subsets.
RegMixMatch optimizes Mixup for semi-supervised learning by integrating high- and low-confidence samples.
WS-II algorithm segments trajectories with high accuracy.
Decision trees can be biased towards minority class, contrary to belief.
Odd-dimensional Riemannian manifolds admit pure spin-c Killing spinors if and only if they are α-Sasakian.
This paper shows hypercommutative algebras on Calabi-Yau manifolds are formal.
We fix integers and . For a -punctured Riemann surface and a -tuple of partitions of , we can define the character variety of type . In this paper, we consider the case where and is indiv…
Non-autoregressive method speeds up protein folding prediction 23 times.
Machine learning models for repeated measurements are limited. Using topological data analysis (TDA), we present a classifier for repeated measurements which samples from the data space and builds a network graph based on the data topology. When applying this to two case studies, accuracy exceeds alternative models wit…
Given a complex projective algebraic variety, write H(X) for its cohomology with complex coefficients and IH(X) for its Intersection cohomology. We first show that, under some fairly general conditions, the canonical map H(X)\to IH(X) is injective. Now let Gr = G((z))/G[[z]] be the loop Grassmannian for a complex semis…
We present a new approach to harmonic analysis that is trained to segment music into a sequence of chord spans tagged with chord labels. Formulated as a semi-Markov Conditional Random Field (semi-CRF), this joint segmentation and labeling approach enables the use of a rich set of segment-level features, such as segment…
We tackle the problem of protein secondary structure prediction using a common task framework. This lead to the introduction of multiple ideas for neural architectures based on state of the art building blocks, used in this task for the first time. We take a principled machine learning approach, which provides genuine,…
Develops a new theory of localization in algebraic geometry.
We describe a strategy for constructing a neural network jet substructure tagger which powerfully discriminates boosted decay signals while remaining largely uncorrelated with the jet mass. This reduces the impact of systematic uncertainties in background modeling while enhancing signal purity, resulting in improved di…
Develops a new theory of localization in algebraic geometry.
Let be a holomorphic vector bundle. Let be a Higgs field, that is a holomorphic section of satisfying . Let be a pluriharmonic metric of the Higgs bundle . The tuple is called a harmonic bundle. Let be a complex manifold, and be a normal crossing divi…
Variational autoencoders are powerful algorithms for identifying dominant latent structure in a single dataset. In many applications, however, we are interested in modeling latent structure and variation that are enriched in a target dataset compared to some background---e.g. enriched in patients compared to the genera…
Convex clustering can only learn convex clusters, with significant gaps between clusters.
Proposes a new clustering method based on expectiles for non-spherical clusters.
CCMM efficiently solves large-scale convex clustering problems.
Discussing issues in robust clustering, especially with Gaussian models.
New indices for determining cluster compactness and separability.
Mode clustering is a nonparametric method for clustering that defines clusters using the basins of attraction of a density estimator's modes. We provide several enhancements to mode clustering: (i) a soft variant of cluster assignment, (ii) a measure of connectivity between clusters, (iii) a technique for choosing the …
EagleEye detects localized density anomalies in multivariate data.
Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of clus…
In this paper, a similarity-driven cluster merging method is proposed for unsuper-vised fuzzy clustering. The cluster merging method is used to resolve the problem of cluster validation. Starting with an overspecified number of clusters in the data, pairs of similar clusters are merged based on the proposed similarity-…
We consider multi-label classification where the goal is to annotate each data point with the most relevant of labels from an extremely large label set. Efficient annotation can be achieved with balanced tree predictors, i.e. trees with logarithmic-depth in the label complexity, whose leaves correspon…
The BPS decomposition theorem splits cohomology of symmetric stacks into invariant parts.
Clustering ensemble, or consensus clustering, has emerged as a powerful tool for improving both the robustness and the stability of results from individual clustering methods. Weighted clustering ensemble arises naturally from clustering ensemble. One of the arguments for weighted clustering ensemble is that elements (…
In many practical applications of clustering, the objects to be clustered evolve over time, and a clustering result is desired at each time step. In such applications, evolutionary clustering typically outperforms traditional static clustering by producing clustering results that reflect long-term trends while being ro…