Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

295887116 · Jun 202019922001200920172026
48 results for cluster purity

We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy c…

2010-10-26abs ↗pdf ↗

Many modern clustering methods scale well to a large number of data items, N, but not to a large number of clusters, K. This paper introduces PERCH, a new non-greedy algorithm for online hierarchical clustering that scales to both massive N and K--a problem setting we term extreme clustering. Our algorithm efficiently …

2017-04-06abs ↗pdf ↗

Graph auto-encoders improve financial clustering using news and stock data.

problem Improving clustering of financial entities using multiple data sources.
method Applying graph deep learning to a finance graph with news co-occurrence and stock price data.
result Dual data sources (news and stock price) improve clustering purity to 64% compared to 32% and 42% for single data sources.

Graph learning categorizes DeFi services into similar functionalities.

problem Identifying similar financial services in decentralized finance protocols.
method Graph representation learning (GRL) to categorize smart contract blocks into clusters.
result Purity of clustering reaches .888 in the best-case scenario.

In this work we propose a simple and easily parallelizable algorithm for multiway graph partitioning. The algorithm alternates between three basic components: diffusing seed vertices over the graph, thresholding the diffused seeds, and then randomly reseeding the thresholded clusters. We demonstrate experimentally that…

2014-06-15abs ↗pdf ↗

In supervised clustering, standard techniques for learning a pairwise dissimilarity function often suffer from a discrepancy between the training and clustering objectives, leading to poor cluster quality. Rectifying this discrepancy necessitates matching the procedure for training the dissimilarity function to the clu…

2019-06-19abs ↗pdf ↗

Clustering analysis by nonnegative low-rank approximations has achieved remarkable progress in the past decade. However, most approximation approaches in this direction are still restricted to matrix factorization. We propose a new low-rank learning method to improve the clustering performance, which is beyond matrix f…

2012-06-18abs ↗pdf ↗

Improves hierarchical clustering in Euclidean space using autoencoders.

problem Lack of unsupervised methods for learning hierarchical structure in Euclidean space.
method Variational autoencoder with Gaussian mixture prior, rescaling latent space, and Ward's linkage.
result Improved dendrogram purity and Moseley-Wang cost function results.

All the connections, pure toward the nilpotent structure, are found. Examples of manifolds, for which the curvature tensor is pure or hybrid, are given. For a manifold of B-type a necessary and sufficient condition for purity of the curvature tensor is proved. It is verified that the conformal change of the metric of a…

2008-07-21abs ↗pdf ↗

This paper examines how regional trade agreements affect global trade relationships.

problem The relationship between regional trade agreements and global trade purity.
method Defined and decomposed synthesized trade resistance, separated natural and artificial factors, used expectation maximization algorithm to optimize parameters, and quantified trade purity indicator.
result Regional trade agreements contribute to the relative prosperity of EU and NAFTA countries, but weaken the role of trade unions and accelerate multilateral trade liberalization.

Point source detection at low signal-to-noise is challenging for astronomical surveys, particularly in radio interferometry images where the noise is correlated. Machine learning is a promising solution, allowing the development of algorithms tailored to specific telescope arrays and science cases. We present DeepSourc…

2018-07-07abs ↗pdf ↗

QNA uses quantum-inspired density operators to diagnose market dependence and structural risk.

problem Lack of unified operator representation for market dependence and structural risk diagnostics.
method Quantum Network of Assets (QNA) framework using density operators.
result QNA entropy remains strongly related to covariance spectral entropy but becomes distinct with multi-feature rolling trajectories.

Deep learning model creates patient representations for scalable EHR-based stratification.

problem Challenges in summarizing and representing patient data from EHRs prevent scalable stratification analysis.
method Unsupervised framework based on deep learning (ConvAE) using word embeddings, CNNs, and autoencoders.
result ConvAE significantly outperformed baselines in clustering diverse patient cohorts, identifying clinically relevant subtypes.

Geometric observables detect financial regime shifts with high accuracy.

problem Detecting regime shifts in financial markets.
method Extracted four geometric observables from equity-index returns and evaluated them against various baseline methods.
result The Berry Phase Rate achieves an unbiased out-of-sample median Cohen's d of 0.72, significantly reducing false alarms.

The study classifies normal subgroups of mapping class groups of surfaces with Cantor subsets.

problem Understanding the structure of normal subgroups in mapping class groups of surfaces with specific subsets.
method Proves two structure theorems: purity and inertia, characterizing normal subgroups.
result Characterizes finite-type normal subgroups of mapping class groups of surfaces with Cantor subsets.

RegMixMatch optimizes Mixup for semi-supervised learning by integrating high- and low-confidence samples.

problem Mixup degrades SSL performance by compromising artificial labels purity.
method RegMixMatch integrates high- and low-confidence samples, uses class-aware Mixup, and mitigates confirmation bias.
result RegMixMatch achieves state-of-the-art performance in SSL benchmarks.

Decision trees can be biased towards minority class, contrary to belief.

problem Bias in decision trees towards minority class in imbalanced datasets.
method Critical evaluation of past literature, specific conditions analysis, tree-fitting adjustments, and post-hoc calibration methods.
result Decision trees can be biased towards minority class under specific conditions, not always towards majority.

We fix integers k>0k> 0 and n>0n>0. For a kk-punctured Riemann surface Σ{p1,,pk}Σ\setminus \{ p_1,\ldots,p_k \} and a kk-tuple μ=(μ1,,μk)\boldsymbolμ=(μ^1,\ldots,μ^k) of partitions of nn, we can define the character variety of type μ\boldsymbolμ. In this paper, we consider the case where Σ=P1Σ=\mathbb{P}^1 and μ\boldsymbolμ is indiv…

2014-06-11abs ↗pdf ↗

Let EE be a holomorphic vector bundle. Let θθ be a Higgs field, that is a holomorphic section of End(E)ΩX1,0End(E)\otimesΩ^{1,0}_X satisfying θ2=0θ^2=0. Let hh be a pluriharmonic metric of the Higgs bundle (E,θ)(E,θ). The tuple (E,θ,h)(E,θ,h) is called a harmonic bundle. Let XX be a complex manifold, and DD be a normal crossing divi…

2002-12-17abs ↗pdf ↗

Variational autoencoders are powerful algorithms for identifying dominant latent structure in a single dataset. In many applications, however, we are interested in modeling latent structure and variation that are enriched in a target dataset compared to some background---e.g. enriched in patients compared to the genera…

2019-02-12abs ↗pdf ↗

Mode clustering is a nonparametric method for clustering that defines clusters using the basins of attraction of a density estimator's modes. We provide several enhancements to mode clustering: (i) a soft variant of cluster assignment, (ii) a measure of connectivity between clusters, (iii) a technique for choosing the …

2014-06-06abs ↗pdf ↗

EagleEye detects localized density anomalies in multivariate data.

problem Identifying signal events, regime changes, or model mismatch in scientific data.
method EagleEye pinpoints local over- and under-densities by assigning anomaly scores based on binary membership sequences and binomial null models.
result EagleEye can detect genuine local anomalies and estimate background purity.

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of clus…

2018-08-24abs ↗pdf ↗

We consider multi-label classification where the goal is to annotate each data point with the most relevant subset\textit{subset} of labels from an extremely large label set. Efficient annotation can be achieved with balanced tree predictors, i.e. trees with logarithmic-depth in the label complexity, whose leaves correspon…

2019-05-24abs ↗pdf ↗

Clustering ensemble, or consensus clustering, has emerged as a powerful tool for improving both the robustness and the stability of results from individual clustering methods. Weighted clustering ensemble arises naturally from clustering ensemble. One of the arguments for weighted clustering ensemble is that elements (…

2019-10-06abs ↗pdf ↗

In many practical applications of clustering, the objects to be clustered evolve over time, and a clustering result is desired at each time step. In such applications, evolutionary clustering typically outperforms traditional static clustering by producing clustering results that reflect long-term trends while being ro…

2011-04-11abs ↗pdf ↗