Proposes σ-PCA to learn identifiable linear transformations without whitening.
problem Cannot identify axes with equal variances in PCA.
method Unified model for linear and nonlinear PCA, introducing a missing piece to eliminate rotational indeterminacy.
result Eliminates subspace rotational indeterminacy in PCA.
KAN-PCA improves asset return analysis by capturing more variance than classical PCA during market crises.
problem Inefficient classical PCA during market crises when correlations between assets change dramatically.
method KAN-PCA uses KAN (Kolmogorov-Arnold Networks) with B-spline functions to learn nonlinear projections.
result KAN-PCA achieves a higher reconstruction R^2 (66.57%) compared to classical PCA (62.99%) on 20 S&P 500 stocks.
We develop efficient algorithms for robust PCA that handle outliers.
problem Finding principal components in datasets with outliers.
method Nearly-linear time and streaming algorithms for robust PCA.
result Near-optimal error guarantees for robust PCA with nearly-linear time and memory usage.
This work connects LLE, factor analysis, and probabilistic PCA through a stochastic perspective.
problem Exploring the theoretical connection between LLE, factor analysis, and probabilistic PCA.
method Solving the stochastic linear reconstruction of LLE using expectation maximization.
result LLE, factor analysis, and probabilistic PCA are shown to be connected through a stochastic perspective.
Simplifies fair PCA with fast, efficient solution.
problem Learning fair low-rank approximations of data.
method Conceptually simple approach with analytic solution.
result Faster and similar results to existing fair PCA methods.
Theoretical validation of linear PCA and ICA for accurate nonlinear BSS.
problem Blind source separation for high-dimensional nonlinear source mixtures.
method Theoretical validation of a cascade of linear PCA and ICA.
result Zero-element-wise-error nonlinear BSS is achieved under certain conditions.
This paper reviews and compares supervised linear dimension-reduction techniques.
problem Lack of information in the response during unsupervised PCA reduces predictive performance.
method Review and comparison of supervised linear dimension-reduction techniques.
result PLS and LSPCA consistently outperform other techniques in simulations.
Wasserstein GAN can approach PCA solution for linear Gaussian data.
problem Understanding GANs' optimality and their relation to PCA.
method Linear-generator Gaussian-data setting, theoretical analysis of GANs and PCA.
result Wasserstein GAN can approach PCA solution in the limit of sample size.
Principal components analysis (PCA) is the optimal linear auto-encoder of data, and it is often used to construct features. Enforcing sparsity on the principal components can promote better generalization, while improving the interpretability of the features. We study the problem of constructing optimal sparse linear a…
PEA improves PCA and k-means for non-linear data and complex clusters.
problem Non-linear dimensionality reduction and clustering challenges.
method Principal Elliptical Analysis (PEA) for efficient non-linear approximation.
result PEA outperforms k-means in complex data clustering.
Diffusion Maps improves on Functional PCA for non-linear functional data.
problem Functional PCA's linear manifold assumption fails for non-linear functional data.
method Extends Diffusion Maps to functional data and compares it to Functional PCA.
result Diffusion Maps outperforms Functional PCA in non-linear functional data analysis.
KPCA improves OoD detection by separating InD and OoD data.
problem Insufficiency of PCA in detecting OoD data from InD data.
method Kernel PCA (KPCA) with task-specific kernels.
result KPCA achieves superior OoD detection performance.
The paper compares PCA and PP for scRNA sequencing data.
problem Limitations of PCA in scRNA sequencing data.
method Applied PCA and PP (using negative Shannon's entropy) on scRNA sequencing data.
result PP outperforms PCA in scRNA sequencing data.
Auto-Associative models cover a large class of methods used in data analysis. In this paper, we describe the generals properties of these models when the projection component is linear and we propose and test an easy to implement Probabilistic Semi-Linear Auto- Associative model in a Gaussian setting. We show it is a g…
This research simplifies PCA model selection using MDL principle.
problem Choosing the right number of principal components in PCA.
method Reduces NML problems to lower-dimension problems and bounds PCA NML.
result Bound the NML of PCA by terms of the NML of linear regression.
One-shot algorithm for feature-distributed kernel PCA reduces communication costs.
problem Efficiently perform kernel PCA in distributed computing environments.
method Inspired by dual relationship between sample-distributed and feature-distributed scenarios, proposes a one-shot algorithm for feature-distributed kernel PCA.
result The algorithm provides high-quality results with low communication costs, especially when eigenvalues decay fast.
New method optimizes PCA for better prediction and variance.
problem Improve PCA for better prediction and variance.
method Jointly optimize prediction error and variance explained.
result Our method outperforms existing approaches in both prediction and variance.
Unified PCA framework on flag manifolds for robust data analysis.
problem Outliers and manifold data in PCA.
method Generalization of PCA to flag manifolds, optimization problems, and tangent-PCA integration.
result Novel robust and dual geodesic PCA variations.
A new dynamical formulation of log-PCA captures local principal modes of geodesic variations.
problem Learning principal variations of random probability measures under Wasserstein geometry.
method Introducing a new dynamical formulation of log-PCA as a variational approach.
result Deriving a general statistical convergence rate for empirical WT-PCA.
Linear principal component analysis (PCA) can be extended to a nonlinear PCA by using artificial neural networks. But the benefit of curved components requires a careful control of the model complexity. Moreover, standard techniques for model selection, including cross-validation and more generally the use of an indepe…
We solve a high-dimensional model where nonlinear autoencoders detect hidden structure missed by PCA.
problem Hidden structure in high-dimensional data not detected by PCA.
method Tractable spiked model with two latent factors, one visible and one uncorrelated.
result Nonlinear autoencoders can extract hidden structure missed by PCA, even if reconstruction loss is higher.
Paper compares PCA of neural network training to high-dimensional random walks.
problem Understanding the dynamics of neural network training through PCA.
method PCA of neural network parameters and random walks, comparison of variances and projections.
result Most variance in neural network training and high-dimensional random walks is captured by the first few PCA components.
SDSPCA improves PCA for disease diagnosis using sparse components and discriminative information.
problem Class ambiguity and low interpretability in traditional PCA.
method Incorporates discriminative information and sparsity into PCA, focusing on sparse components.
result SDSPCA outperforms other methods in gene selection and tumor classification on multi-view biological data.
Adaptive PCA algorithms for changing environments.
problem Static adversarial regret is not suitable for changing environments.
method Online adaptive algorithms for PCA and variance minimization with sub-linear adaptive regret guarantees.
result The proposed algorithms adapt to changing environments.
Sparse PCA selects variables with FDR control for improved performance.
problem Sparse PCA selects irrelevant variables when maximizing explained variance.
method Proposes FDR-controlled selection using T-Rex selector.
result Significant performance improvement over traditional sparse PCA.
Uncertainty-aware PCA preserves data uncertainty during dimensionality reduction.
problem Uncertainty in data affects traditional PCA methods, leading to inaccurate results.
method Generalizes PCA for multivariate probability distributions, respecting uncertainty.
result Uncertainty-aware PCA maintains data characteristics after projection.
Attention learns PCA on Gaussian data, proving its connection to principal component analysis.
problem Principal component analysis on Gaussian data.
method Analysis of attention mechanisms through PCA, covering finite and infinite prompt regimes.
result Attention aligns with principal eigenvectors of covariance matrices, converging to optimal solutions in the infinite-prompt limit.
This paper investigates the generalization of Principal Component Analysis (PCA) to Riemannian manifolds. We first propose a new and general type of family of subspaces in manifolds that we call barycentric subspaces. They are implicitly defined as the locus of points which are weighted means of k+1 reference points.…
A neural network model tackles high-dimensional data with latent structures.
problem Modeling high-dimensional data with latent low-dimensional structures.
method Integrates PCA and Soft PCA layers into neural network architecture for factor modeling and non-linear transformations.
result Demonstrates improved performance in forecasting and nowcasting with real-world data.
We consider the problem of learning a linear factor model. We propose a regularized form of principal component analysis (PCA) and demonstrate through experiments with synthetic and real data the superiority of resulting estimates to those produced by pre-existing factor analysis approaches. We also establish theoretic…
Kernel PCA helps analyze multivariate extremes and clusters them effectively.
problem Analyzing the dependence structure of multivariate extremes.
method Kernel PCA as a method for clustering and dimension reduction.
result Kernel PCA preimages effectively identify clusters in multivariate extremes.
Sparse PCA provides a linear combination of small number of features that maximizes variance across data. Although Sparse PCA has apparent advantages compared to PCA, such as better interpretability, it is generally thought to be computationally much more expensive. In this paper, we demonstrate the surprising fact tha…
Adaptive probabilistic PCA adapts complexity with varying subspaces.
problem Adaptive probabilistic PCA models varying complexity in data.
method Relaxed linear Gaussian model with discrete latent variables, Bayesian nonparametric approach.
result Proposes locally adaptive probabilistic PCA (A-PPCA) for varying subspaces.
Low-precision streaming PCA estimates the leading eigenvector with limited precision.
problem Estimating the leading eigenvector in a streaming setting with limited precision.
method Oja's algorithm with linear and nonlinear stochastic quantization.
result A batched version of the quantized variants achieves the lower bound on quantization error up to logarithmic factors.
Develops an ℓ_p theory for PCA and spectral clustering.
problem Lack of precise characterizations of PCA scores for low-dimensional embedding.
method An ℓ_p perturbation theory for PCA in Hilbert spaces, analyzing eigenvectors and Gram matrix.
result Optimal recovery results for Gaussian mixture and stochastic block models.
Bayesian method improves PCA and MUSIC for unknown number of sources.
problem Estimating the number of sources in PCA and MUSIC algorithms.
method Bayesian inference for computing the exact MAP estimate of the number of sources.
result Bayesian method outperforms AIC in estimating the number of sources.
PCA-based dimensionality reduction improves robustness in overparameterized linear models.
problem Improving robustness in overparameterized linear models.
method PCA-based dimensionality reduction (PCA-OLS)
result PCA-OLS can achieve better generalization than ordinary least squares (OLS) in the overparameterized regime.
The paper analyzes L2-regularized linear autoencoders and their loss landscapes.
problem Understanding the loss landscapes of L2-regularized linear autoencoders. method Smoothly parameterizing the critical manifold and relating minima to the MAP estimate of probabilistic PCA.
result Proves that L2-regularized LAEs learn principal directions as left singular vectors of the decoder. In this paper, a new method is proposed for sparse PCA based on the recursive divide-and-conquer methodology. The main idea is to separate the original sparse PCA problem into a series of much simpler sub-problems, each having a closed-form solution. By recursively solving these sub-problems in an analytical way, an ef…
A new PMA method combines PCA and PLS for better classification.
problem Improving classification performance in data analysis.
method Combines PCA and PLS through multiple sub-PLS models and PCA on the joint coefficient matrix.
result The proposed PMA method achieves better classification performance and stability.
Robust PCA reduces to power iterations for outlier-resilient feature extraction.
problem Sensitivity of PCA to non-Gaussian samples and outliers.
method Robust formulation of PCA based on maximum correntropy criterion.
result MCPI reduces to power iterations, making PCA more robust to outliers.
A new method generalizing subspace learning for improved classification.
problem Improving classification accuracy using subspace learning methods.
method Roweis Discriminant Analysis (RDA) which generalizes PCA, SPCA, and FDA.
result RDA and kernel RDA improve classification accuracy on benchmark datasets.
New method connects Sparse PCA and Sparse Linear Regression.
problem Sparse Principal Component Analysis and Sparse Linear Regression.
method Transforming a solver for Sparse Linear Regression into an algorithm for Sparse Principal Component Analysis.
result The derived SPCA algorithm achieves near state-of-the-art guarantees for testing and support recovery.
Paper proposes PCA-GMM for efficient superresolution of material images.
problem Efficiently handling large and high-dimensional data sets.
method PCA-GMM combining Gaussian Mixture Model and PCA for dimensionality reduction.
result PCA-GMM improves superresolution of material images with moderate dimensionality reduction.
Principal component regression (PCR) is a widely used two-stage procedure: principal component analysis (PCA), followed by regression in which the selected principal components are regarded as new explanatory variables in the model. Note that PCA is based only on the explanatory variables, so the principal components a…
A new method combines predictors and their lags using supervised PCA for dynamic forecasting.
problem Dynamic forecasting with many predictors.
method Supervised PCA with re-scaling and penalized methods.
result The method outperforms traditional PCA and diffusion-index approaches in prediction.
Efficiently approximates Sparse PCA with significant speedups and minor error.
problem Sparse Principal Component Analysis (Sparse PCA) is NP-hard and computationally expensive.
method Approximates the covariance matrix with block-diagonal form, solves sub-problems in each block, and reconstructs the solution.
result Significant computational speedups with minor additive error.
Proposes a new PCA method that balances Euclidean and angle distances.
problem PCA's loss minimization often uses Euclidean distance, but angle distance is more critical in some fields.
method Introduces a method with constraints to unify Euclidean and angle distances, solving the nonconvex optimization problem with an alternating linearized minimization approach.
result Demonstrates the effectiveness and advantages of the new method over state-of-the-art clustering methods on synthetic and real-world datasets.