Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

66132198264 · May 202619922001200920172026
48 results for regime clustering

Study improves S&P 500 volatility forecasting through regime-switching methods.

problem Accurate prediction of S&P 500 volatility for risk management and investment.
method Regime-switching methods including soft Markov switching, spectral clustering, and coefficient-based clustering.
result Coefficient-based clustering algorithm outperformed other models during all time periods.

Clusters asset classes to identify lead-lag relationships in market regimes.

problem Understanding lead-lag relationships between different asset classes.
method Defining macroeconomic regimes by clustering indices and investigating lead-lag relationships.
result Unravels market features and highlights informative market trends or risks.

sWk-means clusters multidimensional financial time series into distinct market regimes.

problem Classifying distinct market regimes in multidimensional financial time series.
method Approximated multidimensional Wasserstein distance as sliced Wasserstein distance for clustering.
result sWk-means successfully identifies distinct market regimes in real financial data.

New method detects and clusters market regimes in multidimensional data.

problem Detecting and clustering market regimes in complex data structures.
method Non-parametric online market regime detection and clustering using path-wise two-sample tests and maximum mean discrepancy.
result Successfully detected and clustered market regimes in various data structures.

The article detects market regimes from covariance matrices using VLSTAR and clustering models.

problem Market regime switching is hard to detect due to time-varying correlation coefficients.
method The article applies VLSTAR and unsupervised hierarchical clustering on monthly realized covariance matrices.
result VLSTAR outperforms clustering in detecting market regimes.

New lower bounds show challenges in clustering in moderate dimensions.

problem Clustering points from mixtures of isotropic Gaussians in moderate dimensions.
method Established low-degree polynomial lower bounds and developed a novel non-spectral algorithm.
result New lower bounds reveal a 'non-parametric rate' in moderate dimensions.

The paper proposes a method to cluster data and estimate regression parameters using VI for financial forecasting.

problem Learning relationships between input and output with different parameters in different regions of the input space.
method Cluster-based regression using Variational Inference (VI).
result The approach can predict the expected value and full distribution of predicted output.

New model clusters mixed-type data with missing values, improving air quality analysis.

problem Clustering mixed-type data with missing values and regime persistence.
method Statistical jump model incorporating regime persistence and handling missing data.
result Superior performance in inferring persistent air quality regimes compared to traditional methods.

For a certain class of distributions, we prove that the linear programming relaxation of kk-medoids clustering---a variant of kk-means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…

2013-09-12abs ↗pdf ↗

A new graph-based clustering method for moderate-dimensional data.

problem Performance degradation of existing graph-based clustering methods in high dimensions.
method Introduces UN-CCDs using NND-based MC-SRT for covering radii determination.
result UN-CCDs provide stable and competitive performance in moderate-sized datasets.

Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.

problem Identifying co-moving assets from correlation matrices for statistical arbitrage.
method Mapping S&P 500 correlation data to GBS-compatible adjacency matrices, benchmarking classical and quantum clustering algorithms.
result Quantum GBS generates superior alpha during high volatility periods, persisting under low-loss conditions.

The paper analyzes kk-means clustering for missing data, proving statistical guarantees under MCAR.

problem Statistical guarantees for kk-means clustering with missing data, especially under Missing Completely at Random (MCAR).
method Established n\sqrt{n}-excess risk bound and consistency of cluster centers under general missing mechanisms; derived n\sqrt{n}-convergence rate and asymptotic normality for MCAR.
result Achieving n\sqrt{n}-rate and converging to true cluster centers requires distinct true cluster centers in every dimension under MCAR.

DIVI clusters noisy high-dimensional data with stable feature gating.

problem Challenging clustering in high-dimensional noisy data.
method Data-informed variational clustering framework combining global feature gating and adaptive structure growth.
result DIVI performs competitively under severe feature noise and remains computationally feasible.

Regularized spectral methods improve clustering in signed graphs, especially for sparse data.

problem Clustering signed graphs with positive and negative edges.
method Developed regularized versions of SPONGE and Signed Laplacian methods for clustering signed graphs, especially for sparse data.
result Theoretical guarantees and empirical performance improvements for clustering signed graphs, especially in sparse regimes.

In this paper, we investigate community detection in networks in the presence of node covariates. In many instances, covariates and networks individually only give a partial view of the cluster structure. One needs to jointly infer the full cluster structure by considering both. In statistics, an emerging body of work …

2016-07-10abs ↗pdf ↗

The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.

problem Ensuring user data privacy in personalized recommendation systems.
method Look-alike clustering to replace sensitive features with cluster averages, analyzed using Convex Gaussian Minimax Theorem.
result Training models using anonymous cluster centers can improve generalization error, especially in high-dimensional settings.

Leveraged ETFs can outperform their targets in certain market conditions, contrary to the volatility drag hypothesis.

problem The long-term performance decay of leveraged ETFs due to volatility drag.
method Unified framework incorporating AR(1) and AR-GARCH models, continuous-time regime switching, and flexible rebalancing frequencies.
result Return dynamics, including return autocorrelation, volatility clustering, and regime persistence, determine LETF performance.

This paper studies the optimality of kernel methods in high-dimensional data clustering. Recent works have studied the large sample performance of kernel clustering in the high-dimensional regime, where Euclidean distance becomes less informative. However, it is unknown whether popular methods, such as kernel k-means, …

2019-12-01abs ↗pdf ↗

Housing markets play a crucial role in economies and the collapse of a real-estate bubble usually destabilizes the financial system and causes economic recessions. We investigate the systemic risk and spatiotemporal dynamics of the US housing market (1975-2011) at the state level based on the Random Matrix Theory (RMT)…

2013-06-12abs ↗pdf ↗

New algorithm for clustering with faulty oracle achieves optimal queries and efficiency.

problem Clustering with a faulty oracle, especially for multiple clusters.
method Built on stochastic block model, provides nearly-optimal query complexity.
result Time-efficient algorithm with nearly-optimal query complexity for all constant k and any δ.

Study clusters distributions with known or unknown clusters using distribution testing.

problem Cluster distributions that are ε\varepsilon-far in total variation.
method Distribution testing approach to establish upper and lower bounds on sample complexity.
result Achieves tight sample complexity bounds for all regimes (up to a logarithmic factor).

Community detection is a fundamental unsupervised learning problem for unlabeled networks which has a broad range of applications. Many community detection algorithms assume that the number of clusters rr is known apriori. In this paper, we propose an approach based on semi-definite relaxations, which does not require…

2017-05-24abs ↗pdf ↗

Quantitative model predicts Sri Lankan stock market using NLP, clustering, and time-series forecasting.

problem Predicting economic regimes and market signals in Sri Lankan stock indices.
method Integrates NLP, clustering, and time-series forecasting; uses FinBERT for sentiment analysis, UMAP/HDBSCAN for clustering, and GRU/LSTM for forecasting.
result GRU model achieves 80.1% R-squared for daily closing price forecasts.

The paper quantizes concatenated noisy vectors to a common cluster center, improving performance over naive methods.

problem Clustering concatenated noisy vectors from multiple sources.
method Asymptotic analysis of weighted sum of distances to a common cluster center.
result The clustering approach outperforms naive methods in terms of average distortion.

The binary symmetric stochastic block model deals with a random graph of nn vertices partitioned into two equal-sized clusters, such that each pair of vertices is connected independently with probability pp within clusters and qq across clusters. In the asymptotic regime of p=alogn/np=a \log n/n and q=blogn/nq=b \log n/n for fixe…

2014-11-24abs ↗pdf ↗

The paper analyzes the dynamics of tokens in transformer models at moderate interaction levels.

problem Understanding the evolution of tokens in transformer models at moderate interaction levels.
method Modeling transformer models as a system of particles interacting in a mean-field way and studying the corresponding dynamics.
result Characterization and convergence of the limiting dynamics in different phases of the system.

This article explores and analyzes the unsupervised clustering of large partially observed graphs. We propose a scalable and provable randomized framework for clustering graphs generated from the stochastic block model. The clustering is first applied to a sub-matrix of the graph's adjacency matrix associated with a re…

2018-05-25abs ↗pdf ↗

We propose infinite mixture prototypes to adaptively represent both simple and complex data distributions for few-shot learning. Our infinite mixture prototypes represent each class by a set of clusters, unlike existing prototypical methods that represent each class by a single cluster. By inferring the number of clust…

2019-02-12abs ↗pdf ↗

This paper examines cryptocurrency integration with traditional markets, showing how network structure and turbulence influence cross-asset spillovers.

problem Understanding how cryptocurrencies integrate with traditional financial markets and the impact of market stress on cross-asset spillovers.
method Combining rolling correlation networks, community structure, market-specific and system-wide Turbulence Indices, and VAR-based connectedness analysis.
result Cross-asset integration is episodic, with network structure and turbulence playing a role in transmission during stress periods.

Hybrid model improves synthetic equity data generation.

problem Generating realistic synthetic financial time series.
method Discretized excess growth rates into states with Poisson jumps, estimating parameters directly.
result Framework achieved high pass rates for distributional and volatility clustering tests.

Spectral clustering performance depends on eigenvector fluctuations, shown to be Gaussian.

problem Predicting the performance of spectral clustering.
method General spike random matrix model and rotational invariance of noise.
result Fluctuations of eigenvector entries are Gaussian in large-dimensional regime.

Bayesian method clusters time series with varying dynamics.

problem Modeling and clustering time series with unknown number of clusters and dynamics.
method Hierarchical Dirichlet process and Gaussian process for modeling time series patterns and variations.
result Efficiently clusters time series with varying dynamics without unnecessary proliferation of clusters.

Spectral clustering is a technique that clusters elements using the top few eigenvectors of their (possibly normalized) similarity matrix. The quality of spectral clustering is closely tied to the convergence properties of these principal eigenvectors. This rate of convergence has been shown to be identical for both th…

2013-10-05abs ↗pdf ↗