PGFL framework learns personalized models with differential privacy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present a framework for clustering with cluster-specific feature selection. The framework, CRAFT, is derived from asymptotic log posterior formulations of nonparametric MAP-based clustering models. CRAFT handles assorted data, i.e., both numeric and categorical data, and the underlying objective functions are intuit…
Mixture of multi-task GPs for clustering and prediction of functional data.
We give an explicit formulaic algorithm and source code for building long-only benchmark portfolios and then using these benchmarks in long-only market outperformance strategies. The benchmarks (or the corresponding betas) do not involve any principal components, nor do they require iterations. Instead, we use a multif…
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
New framework learns complex AI attitudes from heterogeneous data.
Unified framework for clustering and learning causal graphs across subjects.
ARMED models improve deep learning interpretability and generalize better on clustered data.
In this paper, we propose PCKID, a novel, robust, kernel function for spectral clustering, specifically designed to handle incomplete data. By combining posterior distributions of Gaussian Mixture Models for incomplete data on different scales, we are able to learn a kernel for incomplete data that does not depend on a…
The paper proposes a method to cluster data and estimate regression parameters using VI for financial forecasting.
In many applications, multivariate samples may harbor previously unrecognized heterogeneity at the level of conditional independence or network structure. For example, in cancer biology, disease subtypes may differ with respect to subtype-specific interplay between molecular components. Then, both subtype discovery and…
CADM proposes a cluster-specific distance metric for categorical data clustering.
Multi-view clustering is an important approach to analyze multi-view data in an unsupervised way. Among various methods, the multi-view subspace clustering approach has gained increasing attention due to its encouraging performance. Basically, it integrates multi-view information into graphs, which are then fed into sp…
Generative Adversarial networks (GANs) have obtained remarkable success in many unsupervised learning tasks and unarguably, clustering is an important unsupervised learning problem. While one can potentially exploit the latent-space back-projection in GANs to cluster, we demonstrate that the cluster structure is not re…
NeuralFLoC unifies registration and clustering of functional data, overcoming phase variation challenges.
Algorithm tackles clustered contextual bandits with resource constraints.
Feature selection is an important and challenging task in high dimensional clustering. For example, in genomics, there may only be a small number of genes that are differentially expressed, which are informative to the overall clustering structure. Existing feature selection methods, such as Sparse K-means, rarely tack…
This paper presents a neural network-based end-to-end clustering framework. We design a novel strategy to utilize the contrastive criteria for pushing data-forming clusters directly from raw data, in addition to learning a feature embedding suitable for such clustering. The network is trained with weak labels, specific…
New method learns DAG structure in clustered data, accounting for local variations.
We introduce a new, high-throughput, synchronous, distributed, data-parallel, stochastic-gradient-descent learning algorithm. This algorithm uses amortized inference in a compute-cluster-specific, deep, generative, dynamical model to perform joint posterior predictive inference of the mini-batch gradient computation ti…
The smallest eigenvalues and the associated eigenvectors (i.e., eigenpairs) of a graph Laplacian matrix have been widely used for spectral clustering and community detection. However, in real-life applications the number of clusters or communities (say, ) is generally unknown a-priori. Consequently, the majority of …
The smallest eigenvalues and the associated eigenvectors (i.e., eigenpairs) of a graph Laplacian matrix have been widely used in spectral clustering and community detection. However, in real-life applications the number of clusters or communities (say, ) is generally unknown a-priori. Consequently, the majority of t…
MC-GMENN improves neural networks for clustered data using Monte Carlo methods.
Recently there has been an increase in the studies on time-series data mining specifically time-series clustering due to the vast existence of time-series in various domains. The large volume of data in the form of time-series makes it necessary to employ various techniques such as clustering to understand the data and…
Multi-view clustering is a learning paradigm based on multi-view data. Since statistic properties of different views are diverse, even incompatible, few approaches implement multi-view clustering based on the concatenated features straightforward. However, feature concatenation is a natural way to combine multi-view da…
Gradient Boosted Mixed Models estimate mean and variance components for clustered data.
Object clustering, aiming at grouping similar objects into one cluster with an unsupervised strategy, has been extensivelystudied among various data-driven applications. However, most existing state-of-the-art object clustering methods (e.g., single-view or multi-view clustering methods) only explore visual information…
Paper improves short text clustering by integrating semantic relationships into Optimal Transport.
Model predicts higher education dropout risk with interpretable parameters.
A new LDA model with covariates for mixed-membership clusters.
Unified approach for interpretable regression with flexible modeling.
Hierarchical clustering is a class of algorithms that seeks to build a hierarchy of clusters. It has been the dominant approach to constructing embedded classification schemes since it outputs dendrograms, which capture the hierarchical relationship among members at all levels of granularity, simultaneously. Being gree…
Clustering is a fundamental task in data analysis. Recently, deep clustering, which derives inspiration primarily from deep learning approaches, achieves state-of-the-art performance and has attracted considerable attention. Current deep clustering methods usually boost the clustering results by means of the powerful r…
This research improves multitask learning by creating task-specific pathways.
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
Enhances clustering performance with a novel high-order Laplacian matrix.
RR-GNN improves GNN prediction intervals by accounting for graph heteroscedasticity and structural biases.
LLmFPCA-detect detects anomalies in sparse longitudinal text data using LLMs and mFPCA.
Generative model predicts multiple brain graphs from one, preserving topology.
The paper examines partial regularity of Lipschitz solutions to minimal surface system.
Unique ancient solutions found for anisotropic curve shortening flow.
The paper constructs solutions to a critical Dirac equation on spheres.
New findings on -solutions with round cylinder as asymptotic shrinker.
Study higher-dimensional Ricci flow solutions, proving uniqueness.
Ancient solutions of Ricci flow with Type I growth are classified.
We construct low regularity solutions of the vacuum Einstein constraint equations. In particular, on 3-manifolds we obtain solutions with metrics in $H^s\loc$ with . The theory of maximal asymptotically Euclidean solutions of the constraint equations descends completely the low regularity setting. Moreove…
Researchers find entire solutions to magnetic Ginzburg-Landau equations in 4D.
New ancient solutions found for curvature flow in 2D.