Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2.4%4.8%7.1%9.5% · Nov 199519922001200920182026
48 results for Clustering hypothesis

Community detection in networks is a key exploratory tool with applications in a diverse set of areas, ranging from finding communities in social and biological networks to identifying link farms in the World Wide Web. The problem of finding communities or clusters in a network has received much attention from statisti…

2013-11-12abs ↗pdf ↗

Study finds price-based clustering outperforms AI and human methods in stock market analysis.

problem Investigates if AI can improve stock clustering compared to traditional methods.
method Compares price-based, human-informed, and AI-driven clustering methods using synthetic factor models.
result Price-based clustering reduces RMSE by 15.9% relative to GICS and 14.7% relative to LLM embeddings.

Unsupervised clustering is one of the most fundamental challenges in machine learning. A popular hypothesis is that data are generated from a union of low-dimensional nonlinear manifolds; thus an approach to clustering is identifying and separating these manifolds. In this paper, we present a novel approach to solve th…

2017-12-21abs ↗pdf ↗

We describe our language-independent unsupervised word sense induction system. This system only uses topic features to cluster different word senses in their global context topic space. Using unlabeled data, this system trains a latent Dirichlet allocation (LDA) topic model then uses it to infer the topics distribution…

2013-02-28abs ↗pdf ↗

This paper introduces a novel clustering algorithm for heteroscedastic Gaussian data without needing to know the number of clusters.

problem Clustering heteroscedastic Gaussian data without prior knowledge of the number of clusters.
method Introduces a novel cost function and fixed-point analysis to estimate centroids, introduces Wald kernel for measurement plausibility, and derives CENTRE-X algorithm.
result CENTRE-X algorithm can estimate centroids without prior knowledge of the number of clusters and performs comparably to standard algorithms K-means and Mean-Shift.

Cryptocurrencies show varying levels of efficiency over time, forming clusters with younger ones mimicking older ones.

problem Determining the efficiency of cryptocurrencies over time.
method Permutation entropy and statistical complexity over sliding time-windows of price log returns.
result 37% of cryptocurrencies are efficient over 80% of the time, while 20% are efficient in less than 20% of the time.

Network Lasso improves semi-supervised regression on network data.

problem Improving regression accuracy on network data with limited labeled examples.
method Applying network Lasso to semi-supervised regression problems, leveraging message passing over an empirical graph.
result Network Lasso's accuracy is linked to the existence of large network flows over the empirical graph.

A Kyle-inspired model with adaptive agents explains excess volatility and volatility clustering.

problem Reconciling asymmetrically informed traders with adaptive market hypothesis.
method Proposes a model with adaptive agents using inductive reasoning, reconciling Kyle model with Adaptive Market Hypothesis.
result Microfoundations for GARCH models and volatility clustering explained.

A neuro-inspired architecture learns without supervision using clustering and predictive coding.

problem Achieving continual learning without supervision.
method Neuro-inspired architecture based on online clustering and hierarchical predictive coding.
result The architecture achieves continual learning without supervision.

The paper reviews methods for determining the number of communities in network data.

problem Determining the number of communities in network data.
method Statistical methods for hypothesis testing and clustering in network models.
result SCORE and NCV methods evaluated for clustering in Degree-Corrected Block Models, with NCV facing challenges.

The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.

problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.

This paper uses UOT metrics for better dimensionality reduction and classification/clustering.

problem Improving dimensionality reduction and classification/clustering methods.
method Uses Hellinger--Kantorovich metric from unbalanced optimal transport (UOT).
result UOT outperforms Euclidean and OT-based methods in classification and clustering tasks.

This study links blockchain design to cryptos' distributional characteristics.

problem Understanding the relationship between blockchain design and cryptos' distributional characteristics.
method Used spectral clustering to cluster cryptos based on their blockchain mechanisms and operational features.
result Clusters of cryptos share similar blockchain mechanisms, supporting the hypothesis.

Machine learning systems increasingly depend on pipelines of multiple algorithms to provide high quality and well structured predictions. This paper argues interaction effects between clustering and prediction (e.g. classification, regression) algorithms can cause subtle adverse behaviors during cross-validation that m…

2018-07-18abs ↗pdf ↗

The formation of price in a financial market is modelled as a chain of Ising spin with three fundamental figures of trading. We investigate the time behaviour of the model, and we compare the results with the real EURO/USD change rate. By using the test of local Poisson hypothesis, we show that this minimal model leads…

2006-01-09abs ↗pdf ↗

Consider a two-class clustering problem where we observe Xi=iμ+ZiX_i = \ell_i μ+ Z_i, ZiiidN(0,Ip)Z_i \stackrel{iid}{\sim} N(0, I_p), 1in1 \leq i \leq n. The feature vector μRpμ\in R^p is unknown but is presumably sparse. The class labels i{1,1}\ell_i\in\{-1, 1\} are also unknown and the main interest is to estimate them. We are interested …

2015-02-24abs ↗pdf ↗

Paper proposes clustering model for ICC based on histologic patterns.

problem Challenges in grading rare cancers like ICC due to small sample sizes and difficulty in extracting patterns.
method Unsupervised deep convolutional autoencoder clustering model trained on 246 ICC digitized slides.
result Three clusters significantly associated with recurrence-free survival in Cox-proportional hazard models.

The paper develops methods for causal function estimation and inference with multiway clustered data.

problem Estimation and inference for causal functions under multiway clustering.
method Two-step procedure using machine learning for nuisance parameters and projection onto basis functions.
result Rejects the null hypothesis of uniformly zero effects and reveals heterogeneous treatment effects.

A new learning method uses data to learn from large model sets.

problem Learning with large sets of candidate models where uniform convergence is hard.
method Data-dependent learning that incorporates empirical data less reliant on prior assumptions.
result Demonstrates improved generalization in various learning assumptions.

Efficient algorithm for self-directed learning of convex clusters on graphs.

problem Self-directed classification of nodes on graphs with convex clusters.
method Developed efficient algorithms for (geodesically) convex clusters on graphs.
result Polynomial runtime algorithm with 3(h(G)+1)4lnn3(h(G)+1)^4 \ln n mistakes for graphs with two convex clusters.

The Morse-Smale complex of a function ff decomposes the sample space into cells where ff is increasing or decreasing. When applied to nonparametric density estimation and regression, it provides a way to represent, visualize, and compare multivariate functions. In this paper, we present some statistical results on es…

2015-06-29abs ↗pdf ↗

Paper detects and estimates breaks in high-dimensional functional time series.

problem Detecting and estimating structural breaks in heterogeneous mean functions of high-dimensional functional time series.
method Proposes a new test statistic combining functional CUSUM and power enhancement components, with a clustering algorithm for group structure estimation.
result The proposed techniques have satisfactory performance in finite samples, detecting and estimating breaks effectively.

Local elasticity in neural networks makes predictions resilient to dissimilar updates.

problem Understanding resilience of neural network predictions to updates from dissimilar data.
method Simulation and geometric interpretation using neural tangent kernel.
result Local elasticity persists in neural networks with nonlinear activation functions, not in linear ones.

A tick size is the smallest increment of a security price. It is clear that at the shortest time scale on which individual orders are placed the tick size has a major role which affects where limit orders can be placed, the bid-ask spread, etc. This is the realm of market microstructure and there is a vast literature o…

2010-09-13abs ↗pdf ↗

It is widely believed that fluctuations in transaction volume, as reflected in the number of transactions and to a lesser extent their size, are the main cause of clustered volatility. Under this view bursts of rapid or slow price diffusion reflect bursts of frequent or less frequent trading, which cause both clustered…

2005-10-02abs ↗pdf ↗

Method uses TV minimization for semi-supervised learning on network data.

problem Semi-supervised learning from partially-labeled network data.
method Graph signal recovery interpretation, total variation minimization, primal-dual method for non-smooth convex optimization.
result TV minimization recovers clusters in the empirical graph of the data under certain network conditions.

Leveraged ETFs can outperform their targets in certain market conditions, contrary to the volatility drag hypothesis.

problem The long-term performance decay of leveraged ETFs due to volatility drag.
method Unified framework incorporating AR(1) and AR-GARCH models, continuous-time regime switching, and flexible rebalancing frequencies.
result Return dynamics, including return autocorrelation, volatility clustering, and regime persistence, determine LETF performance.

Probabilistic embeddings improve speaker diarization accuracy.

problem Improving speaker diarization accuracy using embeddings.
method Extracting x-vectors and precision matrices from speech segments, interfacing with PLDA model, applying agglomerative clustering, joint training of PLDA and extractor.
result Joint training of PLDA and probabilistic x-vector extractor yields accuracy gains.

Motivated by community detection, we characterise the spectrum of the non-backtracking matrix BB in the Degree-Corrected Stochastic Block Model. Specifically, we consider a random graph on nn vertices partitioned into two equal-sized clusters. The vertices have i.i.d. weights {φu}u=1n\{ φ_u \}_{u=1}^n with second moment $Φ…

2016-09-08abs ↗pdf ↗

A framework for hypothesis testing on attributed graphs using sampling.

problem Statistical testing on graph data, especially large attributed graphs.
method Sampling-based framework with PHASE and PHASEopt for accurate and efficient hypothesis testing.
result PHASE and PHASEopt improve accuracy and efficiency of hypothesis testing in attributed graphs.