HeMPPCAT improves PCA for data with varying noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Dimensionality reduction on Riemannian manifolds is challenging due to the complex nonlinear data structures. While probabilistic principal geodesic analysis~(PPGA) has been proposed to generalize conventional principal component analysis (PCA) onto manifolds, its effectiveness is limited to data with a single modality…
New GMM models fit high-dimensional data with fewer parameters.
This paper improves PPCA robustness using -distributions.
This work connects LLE, factor analysis, and probabilistic PCA through a stochastic perspective.
Paper proposes PCA-GMM for efficient superresolution of material images.
Probabilistic Autoencoder learns latent space weights' distribution.
Principal component analysis (PCA) is arguably the most popular tool in multivariate exploratory data analysis. In this paper, we consider the question of how to handle heterogeneous variables that include continuous, binary, and ordinal. In the probabilistic interpretation of low-rank PCA, the data has a normal multiv…
Principal Component Analysis (PCA) is a popular tool for dimensionality reduction and feature extraction in data analysis. There is a probabilistic version of PCA, known as Probabilistic PCA (PPCA). However, standard PCA and PPCA are not robust, as they are sensitive to outliers. To alleviate this problem, this paper i…
Develops an ℓ_p theory for PCA and spectral clustering.
A method to remove mean-shift noise from PCA using knockoffs.
Generative models' ELBOs converge to entropy sums, proving for various models.
Paper develops a dual formulation for PCA in Hilbert spaces.
We discuss a clustering method for Gaussian mixture model based on the sparse principal component analysis (SPCA) method and compare it with the IF-PCA method. We also discuss the dependent case where the covariance matrix is not necessarily diagonal.
Sparse versions of principal component analysis (PCA) have imposed themselves as simple, yet powerful ways of selecting relevant features of high-dimensional data in an unsupervised manner. However, when several sparse principal components are computed, the interpretation of the selected variables is difficult since ea…
We consider the problem of learning a mixture of Random Utility Models (RUMs). Despite the success of RUMs in various domains and the versatility of mixture RUMs to capture the heterogeneity in preferences, there has been only limited progress in learning a mixture of RUMs from partial data such as pairwise comparisons…
Paper revisits PCA for anomaly detection in network security.
Transformers interpret as probabilistic mixtures, offering new insights.
Principal component analysis (PCA) is one of the most widely used dimension reduction and multivariate statistical techniques. From a probabilistic perspective, PCA seeks a low-dimensional representation of data in the presence of independent identical Gaussian noise. Probabilistic PCA (PPCA) and its variants have been…
Auto-Associative models cover a large class of methods used in data analysis. In this paper, we describe the generals properties of these models when the projection component is linear and we propose and test an easy to implement Probabilistic Semi-Linear Auto- Associative model in a Gaussian setting. We show it is a g…
Mixtures of Gaussians, factor analyzers (probabilistic PCA) and hidden Markov models are staples of static and dynamic data modeling and image and video modeling in particular. We show how topographic transformations in the input, such as translation and shearing in images, can be accounted for in these models by inclu…
We present a method to compute the Shapley values of reconstruction errors of principal component analysis (PCA), which is particularly useful in explaining the results of anomaly detection based on PCA. Because features are usually correlated when PCA-based anomaly detection is applied, care must be taken in computing…
We consider probabilistic PCA and related factor models from a Bayesian perspective. These models are in general not identifiable as the likelihood has a rotational symmetry. This gives rise to complicated posterior distributions with continuous subspaces of equal density and thus hinders efficiency of inference as wel…
Hybrid model combines continuous and tractable probabilistic models.
Survey of factor analysis, PCA, variational inference, and VAE.
CAVI converges exponentially fast for Bayesian PCA models.
Paper proposes GPM for heteroscedastic PCA estimation.
PGPCA improves PCA for nonlinear data in neuroscience.
New insights into risk aversion for complex decision models.
We shed new insights on the two commonly used updates for the online -PCA problem, namely, Krasulina's and Oja's updates. We show that Krasulina's update corresponds to a projected gradient descent step on the Stiefel manifold of the orthonormal -frames, while Oja's update amounts to a gradient descent step using…
Proposes MPCA for robust PCA using mode estimation.
uGMM-NN integrates probabilistic reasoning into neural networks.
Proposes deep mixture models for probabilistic price movement forecasting in high-frequency trading.
In this paper, we study the spectrum and the eigenvectors of radial kernels for mixtures of distributions in . Our approach focuses on high dimensions and relies solely on the concentration properties of the components in the mixture. We give several results describing of the structure of kernel matrices …
The paper presents a method to reduce computational and storage costs in PCA and spectral clustering.
Combines PCA and CEM for fast clustering and embedding.
For many years, a combination of principal component analysis (PCA) and independent component analysis (ICA) has been used for blind source separation (BSS). However, it remains unclear why these linear methods work well with real-world data that involve nonlinear source mixtures. This work theoretically validates that…
We introduce Probabilistic FastText, a new model for word embeddings that can capture multiple word senses, sub-word structure, and uncertainty information. In particular, we represent each word with a Gaussian mixture density, where the mean of a mixture component is given by the sum of n-grams. This representation al…
Paper uses GMM and MAF for probabilistic classification, outperforming simpler models.
A new approach to continual learning using fully probabilistic models.
Small neural networks embed arbitrary metric spaces into Gaussian mixtures.
MD-CGAN models forecast time series with probabilistic posterior distributions.
A new method for sparse PCA using orthogonal rotations and soft-thresholding.
Evolutionary clustering aims at capturing the temporal evolution of clusters. This issue is particularly important in the context of social media data that are naturally temporally driven. In this paper, we propose a new probabilistic model-based evolutionary clustering technique. The Temporal Multinomial Mixture (TMM)…
A new probabilistic polygonal curve representation using Gaussian Mixture Models.
We present a unifying framework which reduces the construction of probabilistic component analysis techniques to a mere selection of the latent neighbourhood, thus providing an elegant and principled framework for creating novel component analysis models as well as constructing probabilistic equivalents of deterministi…
The book explores universal time-series forecasting using mixture predictors.
Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…