Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4691137182 · Jun 202019922001200920172026
48 results for Universal Clustering

A novel method relaxes binary constraints to non-negative spheres for multi-matching and clustering.

problem Optimization problems over binary matrices with injectivity constraints.
method Non-negative spherical relaxation followed by conditional power iteration.
result Automatic adjustment of the continuous parameter related to universe size.

Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.

problem Achieving optimal error rates in clustering sub-exponential mixture models.
method Establishes universal lower bounds and demonstrates iterative algorithms' optimality in sub-exponential mixture models.
result Iterative algorithms achieve the universal lower bound in sub-exponential mixture models.

Clustering is one of the most universal approaches for understanding complex data. A pivotal aspect of clustering analysis is quantitatively comparing clusterings; clustering comparison is the basis for many tasks such as clustering evaluation, consensus clustering, and tracking the temporal evolution of clusters. In p…

2017-06-19abs ↗pdf ↗

NOs can learn any finite collection of classes in functional data.

problem Learning finite collections of classes in infinite-dimensional spaces.
method Proved sample-based neural operators can learn any finite collection of classes in an infinite-dimensional reproducing kernel Hilbert space.
result NOs can learn any finite collection of classes in an infinite-dimensional reproducing kernel Hilbert space, even when the classes are not convex or connected.

New criterion selects optimal number of clusters based on stability.

problem Challenges in selecting optimal number of clusters in non-parametric clustering.
method Proposes a stability-based validation criterion combining between-cluster and within-cluster stability.
result Empirically demonstrates effectiveness in selecting optimal number of clusters.

Subspace clustering refers to the problem of clustering high-dimensional data into a union of low-dimensional subspaces. Current subspace clustering approaches are usually based on a two-stage framework. In the first stage, an affinity matrix is generated from data. In the second one, spectral clustering is applied on …

2019-10-20abs ↗pdf ↗

Consider unsupervised clustering of objects drawn from a discrete set, through the use of human intelligence available in crowdsourcing platforms. This paper defines and studies the problem of universal clustering using responses of crowd workers, without knowledge of worker reliability or task difficulty. We model sto…

2016-10-05abs ↗pdf ↗

TACE unifies scalar and tensorial modeling in Cartesian space for accurate, stable, and efficient atomistic predictions.

problem Complexity and challenges in equivariant atomistic machine learning models.
method Tensor Atomic Cluster Expansion (TACE) in Cartesian space, decomposing local environments into irreducible Cartesian tensors (ICT).
result Universal invariant and equivariant embeddings, enabling explicit control at inference.

Method detects lead-lag relationships in multivariate time series.

problem Discovering lead-lag relationships in multivariate time series.
method Clustering-driven methodology using sliding window and various clustering techniques.
result Robust lead-lag estimates across clusters enhance consistent relationships identification.

The paper uses clustering and integer programming to optimize stock selection for investment funds.

problem Maximizing profits and minimizing risk in stock markets.
method Data-oriented analysis and clustering techniques with integer programming.
result Reconstructed NASDAQ 100 index fund example demonstrates effectiveness.

Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often makes the algorithm selection and hyperparameter evaluation a tough guess. In t…

2018-03-29abs ↗pdf ↗

NeuralFLoC unifies registration and clustering of functional data, overcoming phase variation challenges.

problem Challenges in clustering functional data due to phase variation and temporal misalignment.
method NeuralFLoC uses Neural ODE-driven diffeomorphic flows and spectral clustering for joint registration and clustering.
result NeuralFLoC effectively disentangles phase and amplitude variation, achieving state-of-the-art performance.

We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…

2017-03-02abs ↗pdf ↗

Financial price changes obey two universal properties: they follow a power law and they tend to be clustered in time. The second regularity, known as volatility clustering, entails some predictability in the price changes: while their sign is uncorrelated in time, their amplitude (or volatility) is long-range correlate…

2016-12-29abs ↗pdf ↗

New method clusters directed graphs using Koopman operators.

problem Challenges in clustering directed graphs, especially complex eigenvalues and lack of cluster definition.
method Relate graph Laplacians to transfer operators and metastable sets in stochastic systems, derive clustering algorithms for directed and time-evolving graphs.
result Clusters can be interpreted as coherent sets, useful for analyzing transport and mixing processes.

Clusters of crypto assets by path signature improve diversification and reduce fees.

problem Building diversified portfolios of volatile cryptocurrencies.
method Clustering digital assets using path signatures to identify similar behavior patterns.
result Optimal portfolios outperform unfiltered ones, reducing transaction fees.

Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.

problem Predicting the performance of spectral clustering.
method General spike random matrix model and rotational invariance of noise.
result Fluctuations of eigenvector entries are Gaussian in large-dimensional regime.

In this paper, we present a novel way to summarize the structure of large graphs, based on non-parametric estimation of edge density in directed multigraphs. Following coclustering approach, we use a clustering of the vertices, with a piecewise constant estimation of the density of the edges across the clusters, and ad…

2015-08-06abs ↗pdf ↗

We propose an efficient Context-Aware clustering of Bandits (CAB) algorithm, which can capture collaborative effects. CAB can be easily deployed in a real-world recommendation system, where multi-armed bandits have been shown to perform well in particular with respect to the cold-start problem. CAB utilizes a context-a…

2015-10-12abs ↗pdf ↗

Model predicts higher education dropout risk with interpretable parameters.

problem Predicting and understanding student dropout risk in higher education.
method Sparse interpretable post-clustering logistic regression.
result Model identifies distinct dropout risk subgroups within the student population.

MPE framework proves universal approximation for quantum data distribution.

problem Challenges in generating quantum data from underlying distributions.
method Many-body Projected Ensemble (MPE) framework for quantum state design.
result MPE can approximate any quantum distribution within 1-Wasserstein distance error.

The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.

problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.

Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.

problem Identifying co-moving assets from correlation matrices for statistical arbitrage.
method Mapping S&P 500 correlation data to GBS-compatible adjacency matrices, benchmarking classical and quantum clustering algorithms.
result Quantum GBS generates superior alpha during high volatility periods, persisting under low-loss conditions.

Market sectors play a key role in the efficient flow of capital through the modern Global economy. We analyze existing sectorization heuristics, and observe that the most popular - the GICS (which informs the S&P 500), and the NAICS (published by the U.S. Government) - are not entirely quantitatively driven, but rather…

2019-05-31abs ↗pdf ↗

The study examines the universality of Gaussian data in high-dimensional generalized linear estimation.

problem Understanding when Gaussian data suffices for high-dimensional generalized linear estimation.
method Sharp asymptotic expressions for test and training errors in high-dimensional Gaussian mixture data with labels from a single-index model.
result The universality of Gaussian data in error estimation depends on the alignment between target weights and mixture cluster means and covariances.

Quantum cluster algebra constructed from web skein relations on surfaces.

problem Quantization of cluster structures on moduli spaces of SL3 local systems.
method Constructing a quantum cluster algebra inside the skew-field of a skein algebra of unpunctured surfaces.
result Laurent expressions of webs in clusters have positive coefficients.

Transformers learn to cluster Gaussian mixtures as well as the EM algorithm.

problem Learning guarantees of Transformers in multi-class clustering of Gaussian mixtures.
method Developed a theory connecting Transformer's Softmax Attention layers to the EM algorithm's workflow.
result Transformers achieve minimax optimal rate for clustering Gaussian mixtures with sufficient training samples and initialization.

We introduce the cluster exchange groupoid associated to a non-degenerate quiver with potential, as an enhancement of the cluster exchange graph. In the case that arises from an (unpunctured) marked surface, where the exchange graph is modelled on the graph of triangulations of the marked surface, we show that the univ…

2018-04-30abs ↗pdf ↗

Study of 2D Ising model reveals patterns in financial markets.

problem Understanding stylized facts in financial markets using statistical physics.
method 2D Ising model with spin interactions; analysis of spin clusters, persistence, and dynamics.
result Microscopic mechanisms explain stylized facts like sharp peaks in returns and heavy-tailed distributions.

Unified HDP and LDA models for efficient topic clustering of online course queries.

problem Efficiently cluster and answer subject-specific online course queries.
method Use Hierarchical Dirichlet Process (HDP) to optimize topic number for Latent Dirichlet Allocation (LDA) model runs.
result Achieve optimal clustering efficiency by recursively applying LDA on effective topics.

Modeling videos and image-sets as linear subspaces has proven beneficial for many visual recognition tasks. However, it also incurs challenges arising from the fact that linear subspaces do not obey Euclidean geometry, but lie on a special type of Riemannian manifolds known as Grassmannian. To leverage the techniques d…

2014-07-04abs ↗pdf ↗