Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

4590135180 · Jun 202019922001200920172026
48 results for covariate clustering

A new method clusters covariates considering class labels for better classification.

problem Clustering covariates independently of class labels can lead to poor results.
method Formulates as convex optimization, uses ADMM for solving, and selects model via marginal likelihood.
result Proposed method offers a unique global minimum and improves classification.

Biological and social systems consist of myriad interacting units. The interactions can be represented in the form of a graph or network. Measurements of these graphs can reveal the underlying structure of these interactions, which provides insight into the systems that generated the graphs. Moreover, in applications s…

2014-11-08abs ↗pdf ↗

New method clusters high-dimensional data with anisotropic noise.

problem Clustering high-dimensional anisotropic mixtures with varying noise structures.
method Covariance Projected Spectral Clustering (COPO) method that projects data onto a low-dimensional space and reassigns clusters based on estimated covariances.
result COPO achieves minimax-optimal misclustering rates in Gaussian settings.

Develops a new random forest method for clustered data with improved prediction and inference.

problem Improving prediction and inference accuracy for clustered data with within-cluster dependence.
method Clustered Random Forests, using weighted least squares estimators for leaf predictions.
result Optimal prediction and inference weights vary under covariate shift, necessitating user-chosen weights.

In this paper, we investigate community detection in networks in the presence of node covariates. In many instances, covariates and networks individually only give a partial view of the cluster structure. One needs to jointly infer the full cluster structure by considering both. In statistics, an emerging body of work …

2016-07-10abs ↗pdf ↗

A new LDA model with covariates for mixed-membership clusters.

problem Modeling mixed-membership clusters in discrete data with covariates.
method Negative binomial regression embedded within LDA, slice sampling within Gibbs sampling.
result Model successfully retrieves true parameter values and predicts cluster abundances using covariates.

Spatially relaxed inference tackles high-dimensional linear models with correlated covariates.

problem Accurate inference is challenging in high-dimensional settings with spatially correlated covariates.
method Proposes ensembled clustered inference algorithms that control the δδ-FWER under standard assumptions.
result Ensembled clustered inference algorithms control the δδ-FWER and achieve decent power.

This study evaluates cluster search algorithms using Gaussian mixture models.

problem Determining the optimal number of clusters in data sets generated by Gaussian mixture models.
method Examined centroid- and model-based cluster search algorithms in various cases.
result Model-based algorithms are more robust to cluster overlap and covariance type than centroid-based methods.

GBMixed boosts mixed models for clustered data, estimating mean and variance flexibly.

problem Flexible estimation of mean and variance components in clustered data.
method Gradient Boosting framework for linear mixed models with likelihood-based gradients.
result GBMixed accurately recovers complex nonlinear fixed effects and covariances.

Algorithm clusters mixtures with bounded covariances under specific separation conditions.

problem Clustering mixtures of bounded covariance distributions with fine-grained separation.
method Introduced clustering refinement and efficient algorithm for accurate clustering.
result First poly-time algorithm for nearly uniform mixtures, and efficient refinement for general mixtures.

A new algorithm COVA-FC improves subgroup-fair clustering efficiency.

problem Challenges in making cluster assignments independent of sensitive attributes in subgroups.
method Defining a subgroup-fairness gap, deriving a covariance-based surrogate, and introducing a continuous relaxation for efficient optimization.
result COVA-FC achieves competitive cost-fairness trade-offs and improves computational efficiency.

Proposes ICC method for dynamic portfolio optimization.

problem Non-stationarity in market conditions makes traditional portfolio optimization ineffective.
method Inverse Covariance Clustering (ICC) to identify market states and integrate into dynamic optimization.
result ICC-PO generates portfolios with higher Sharpe Ratios and greater robustness.

Regularized EM algorithm improves clustering performance with small sample sizes.

problem Performance reduction in EM algorithm due to small sample size and poorly conditioned covariance matrices.
method Regularized EM algorithm that uses prior knowledge to ensure positive definiteness of covariance matrices.
result The regularized EM algorithm outperforms standard EM in clustering tasks with small sample sizes.

Study improves portfolio risk estimation methods using robust covariance and CVaR constraints.

problem Improving portfolio risk estimation in the presence of financial data noise and extreme market conditions.
method Exploration of robust covariance estimators, application of CVaR constraints, use of K-means clustering in optimization.
result Robust covariance estimators can outperform market-weighted benchmarks, especially during bull markets.

Efficiently clusters nodes in Gaussian graphical models from data.

problem Clustering nodes in Gaussian graphical models directly from data.
method Clusters nodes based on the similarity of their network neighborhoods defined by partial correlations. Uses matrix factors for limited data.
result Demonstrates improved clustering of nodes in Gaussian graphical models.

CDL index improves clustering validation for non-convex data.

problem Selecting clustering algorithms and hyperparameters without labeled data.
method CDL uses compactness, centers, and covariances to compute a probabilistic description length bound.
result CDL outperforms conventional CVIs on synthetic and image benchmarks.

We consider the problem of analyzing the heterogeneity of clustering distributions for multiple groups of observed data, each of which is indexed by a covariate value, and inferring global clusters arising from observations aggregated over the covariate domain. We propose a novel Bayesian nonparametric method reposing …

2010-01-04abs ↗pdf ↗

Regularized EM algorithm improves GMM clustering in low sample settings.

problem Numerical instability and convergence issues in EM-GMM for low sample support.
method Regularized EM algorithm that maximizes penalized GMM likelihood, ensuring positive definiteness and structured covariance matrices.
result The regularized EM algorithm leads to better performing EM for structured covariance matrix models or low sample settings.

Co-trading networks reveal dynamic market structures and improve covariance estimation.

problem Modeling high-dimensional stock covariances in US equity markets.
method Co-trading-based pairwise similarity measure for constructing dynamic networks, spectral clustering, robust covariance estimator.
result Co-trading networks capture time-evolving stock dependencies and improve portfolio performance.

Asymptotically consistent clustering algorithms for ergodic stochastic processes are developed.

problem Clustering stochastic processes with consistency guarantees.
method Review and development of clustering algorithms for ergodic stochastic processes.
result Asymptotically consistent clustering algorithms can be obtained for ergodic stochastic processes.

New algorithm speeds up cluster-based compressive sensing tasks.

problem Efficiently solving multiple compressive sensing tasks with shared information.
method Combines Monte Carlo sampling with iterative linear solvers to avoid explicit covariance matrix computation.
result Up to thousands of times faster and orders of magnitude more memory-efficient compared to existing methods.

New method for estimating financial covariance matrices efficiently.

problem Noisy covariance matrix estimation in high-dimensional financial data.
method Cluster financial time series into groups, apply shrinkage to ensure positive definiteness.
result Proposed methods provide reliable estimates and outperform other estimators.

This paper introduces a novel clustering algorithm for heteroscedastic Gaussian data without needing to know the number of clusters.

problem Clustering heteroscedastic Gaussian data without prior knowledge of the number of clusters.
method Introduces a novel cost function and fixed-point analysis to estimate centroids, introduces Wald kernel for measurement plausibility, and derives CENTRE-X algorithm.
result CENTRE-X algorithm can estimate centroids without prior knowledge of the number of clusters and performs comparably to standard algorithms K-means and Mean-Shift.

Study on clustering in high dimensions with anisotropic Gaussian mixtures, showing interpolation can be optimal and robust.

problem Clustering in high-dimensional anisotropic Gaussian mixtures.
method Derive minimax bounds, analyze 2\ell_2-regularized classifiers, and investigate interpolation's robustness.
result Interpolating solutions can be optimal and robust under certain conditions.

We discuss a clustering method for Gaussian mixture model based on the sparse principal component analysis (SPCA) method and compare it with the IF-PCA method. We also discuss the dependent case where the covariance matrix ΣΣ is not necessarily diagonal.

2016-02-16abs ↗pdf ↗

MCAP clusters high-dimensional data via adaptive projections, handling large p efficiently.

problem Statistical and computational challenges in high-dimensional mixture models.
method Model-based Clustering via Adaptive Projections (MCAP) using linear projections.
result MCAP reliably detects covariance signals in very high-dimensional problems.

The article detects market regimes from covariance matrices using VLSTAR and clustering models.

problem Market regime switching is hard to detect due to time-varying correlation coefficients.
method The article applies VLSTAR and unsupervised hierarchical clustering on monthly realized covariance matrices.
result VLSTAR outperforms clustering in detecting market regimes.

The paper proposes a new portfolio allocation method combining RMT and machine learning.

problem Optimal allocation instability in high-dimensional portfolios.
method Combines Random Matrix Theory covariance estimators with Nested Clustered Optimization.
result The modified NCO algorithm achieves stable allocations without risky short positions.

Two-stage TMLE reduces bias and improves efficiency in CRTs.

problem Differential outcome measurement and imbalance in baseline predictors in CRTs.
method Two-stage targeted minimum loss-based estimator (TMLE) to adjust for baseline covariates.
result Our approach nearly eliminates bias due to differential outcome measurement.

Study spectral properties of radial kernels for high-dimensional mixtures.

problem Understanding spectral properties of radial kernels for high-dimensional mixtures.
method High-dimensional analysis focusing on concentration properties of components in mixtures.
result Kernel PCA can successfully cluster mixtures with common means but different covariances, even in high dimensions.