New method corrects missing data bias in dimension reduction.
problem Missing data complicates high-dimensional data analysis.
method Developed a bias-corrected Gram matrix for heterogeneous missingness.
result Proposed method improves dimension reduction techniques significantly.
Survey on geometric foundations of data reduction methods.
problem High-dimensional data with intrinsic nonlinear structure.
method Spectral manifold learning methods.
result Derivation and convergence analysis of spectral manifold learning.
Consensus dimension reduction combines multiple visualizations to identify shared patterns.
problem Conflicting visualizations from different dimension reduction methods.
method Multi-view learning to identify stable patterns across multiple views.
result Consensus visualization effectively identifies shared low-dimensional data structure.
Data balancing reduces variance in machine learning models.
problem Reduction of variance in machine learning models.
method Non-asymptotic statistical bound and eigenvalue decay of Markov operators.
result Data balancing across modalities and sources reduces variance.
New neural network method simplifies high-dimensional data.
problem Scalability issues in nonlinear sufficient dimension reduction.
method Stochastic neural network with adaptive gradient algorithm.
result Proposed method outperforms existing methods on large-scale data.
In statistical learning, high covariate dimensionality poses challenges for robust prediction and inference. To address this challenge, supervised dimension reduction is often performed, where dependence on the outcome is maximized for a selected covariate subspace with smaller dimensionality. Prevalent dimension reduc…
Enhances SDR via Hellinger correlation for better data dependency understanding.
problem Improving sufficient dimension reduction in single-index models.
method Developed a new method using Hellinger correlation for detecting the dimension reduction subspace.
result Significantly enhances and outperforms existing SDR methods through deeper data dependency understanding.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
New DR method uses Gromov-Wasserstein distance for high-dimensional data.
problem Analyzing relationships between high-dimensional objects.
method Optimal transportation theory and Gromov-Wasserstein distance.
result Robust and efficient solution for complex high-dimensional datasets.
Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain …
Scalability of statistical estimators is of increasing importance in modern applications and dimension reduction is often used to extract relevant information from data. A variety of popular dimension reduction approaches can be framed as symmetric generalized eigendecomposition problems. In this paper we outline how t…
Modern techniques simplify complex high-dimensional data.
problem Complex, high-dimensional data.
method Unsupervised dimension reduction techniques.
result Simplified representation of high-dimensional data.
t-SNE loses important features in data visualization.
problem t-SNE's loss of important features in data visualization.
method Established mathematical framework to understand t-SNE's loss in different scenarios.
result t-SNE loses important features of data in various scenarios.
Modeling data as being sampled from a union of independent subspaces has been widely applied to a number of real world applications. However, dimensionality reduction approaches that theoretically preserve this independence assumption have not been well studied. Our key contribution is to show that 2K projection vect…
New algorithm balances spatial data approximation and prediction accuracy.
problem Lack of methods considering spatial correlation and downstream modeling in dimension reduction.
method Formalizes approximation and modeling utility as metrics, proposes a balanced algorithm.
result Optimal trade-off between approximation accuracy and downstream modeling utility.
Unified framework for DR and clustering using Gromov-Wasserstein.
problem Capturing structure in high-dimensional datasets.
method Distributional reduction framework using Gromov-Wasserstein.
result Unified approach recovers DR and clustering as special cases.
This paper addresses overfitting in dimension reduction methods by calibrating hyperparameters considering noise.
problem Overfitting in dimension reduction methods, especially t-SNE and UMAP, when data contains noise.
method Present a framework to calibrate hyperparameters in the presence of noise for t-SNE and UMAP.
result Recommended hyperparameter values for t-SNE and UMAP are too small and overfit the noise.
This work compares data reduction criteria for online Gaussian Processes.
problem The computational complexity of Gaussian Processes limits their applicability to small datasets and streaming scenarios.
method Unified comparison of several data reduction criteria, analyzing computational complexity and reduction behavior.
result Practical guidelines for choosing a suitable data reduction criterion for online Gaussian Processes.
POTD estimates SDR subspace using optimal transport for binary response.
problem Insufficient performance of existing SDR methods for categorical responses.
method Principal optimal transport direction (POTD) using optimal transport coupling.
result POTD exclusively estimates SDR subspace for error-free class labels.
We survey the role of symmetry in diffeomorphic registration of landmarks, curves, surfaces, images and higher-order data. The infinite dimensional problem of finding correspondences between objects can for a range of concrete data types be reduced resulting in compact representations of shape and spatial structure. Th…
In this work, we revisit fast dimension reduction approaches, as with random projections and random sampling. Our goal is to summarize the data to decrease computational costs and memory footprint of subsequent analysis. Such dimension reduction can be very efficient when the signals of interest have a strong structure…
Class prediction is an important application of microarray gene expression data analysis. The high-dimensionality of microarray data, where number of genes (variables) is very large compared to the number of samples (obser- vations), makes the application of many prediction techniques (e.g., logistic regression, discri…
We present local discriminative Gaussian (LDG) dimensionality reduction, a supervised dimensionality reduction technique for classification. The LDG objective function is an approximation to the leave-one-out training error of a local quadratic discriminant analysis classifier, and thus acts locally to each training po…
In this era of data deluge, many signal processing and machine learning tasks are faced with high-dimensional datasets, including images, videos, as well as time series generated from social, commercial and brain network interactions. Their efficient processing calls for dimensionality reduction techniques capable of p…
Nonlinear dimensionality reduction methods are a popular tool for data scientists and researchers to visualize complex, high dimensional data. However, while these methods continue to improve and grow in number, it is often difficult to evaluate the quality of a visualization due to a variety of factors such as lack of…
Paper compares dimension reduction methods using topological analysis on EEG data.
problem Comparing dimension reduction methods on EEG data.
method Topological data analysis, including persistent homology, Wasserstein distance, and hypothesis tests.
result Different dimension reduction methods show significant qualitative differences across topological homologies.
Develops a new method for nonlinear dimension reduction using random features.
problem Statistical challenges in generalizing Gaussian process-based latent variable models to non-Gaussian data.
method Random feature latent variable models (RFLVMs) that approximate nonlinear relationships with linear functions of random features.
result RFLVMs produce comparable results to state-of-the-art methods on various data types.
Adapts DR objectives for both sample and feature size reduction.
problem Simultaneously reduce sample and feature sizes.
method Semi-relaxed Gromov-Wasserstein optimal transport.
result OT plan delivers competitive hard clustering.
Paper proposes a tensor data model for incomplete imaging data.
problem Prognostics models for incomplete imaging data.
method Supervised tensor dimension reduction with TTF supervision and optimization.
result Model effectively extracts low-dimensional features from incomplete data.
New algorithms for clustering and dimension reduction using relative von Neumann entropy.
problem Clustering and dimension reduction for complex data sets.
method Construct graphs from data points, select graph maximizing relative von Neumann entropy, use eigenvectors for dimension reduction.
result Outperforms existing methods on non-trivial data sets.
A new framework learns clustering and dimensionality reduction together.
problem Challenges in clustering high-dimensional data.
method Gradient-based manifold optimization for joint learning.
result Better performance compared to existing clustering algorithms.
Experimental life sciences like biology or chemistry have seen in the recent decades an explosion of the data available from experiments. Laboratory instruments become more and more complex and report hundreds or thousands measurements for a single experiment and therefore the statistical methods face challenging tasks…
A new geometry-preserving method for interpreting compositional data.
problem Statistical challenges in high-dimensional compositional data.
method Geometry-preserving framework for dimension reduction of compositional data.
result Identification of a central compositional subspace for compositional predictors.
WeldNet reduces complex dynamics to simpler, manageable segments.
problem Complex, high-dimensional time-dependent datasets from physical processes are costly to simulate.
method Windowed Encoders for Learning Dynamics, splitting time domain into windows for nonlinear dimension reduction and propagator training.
result WeldNet captures nonlinear latent structures and dynamics, outperforming existing methods.
Unified model for reducing dimensions and clustering high-dimensional data.
problem High-dimensional data clustering and dimensionality reduction.
method Hierarchical mixtures of Gaussians (HMoGs) with closed-form likelihood and inference.
result Efficiently models hundreds of latent dimensions, improving clustering performance.
In this paper, we study randomized reduction methods, which reduce high-dimensional features into low-dimensional space by randomized methods (e.g., random projection, random hashing), for large-scale high-dimensional classification. Previous theoretical results on randomized reduction methods hinge on strong assumptio…
PGPCA improves PCA for nonlinear data in neuroscience.
problem Nonlinear data distribution in neuroscience.
method Developed PGPCA for nonlinear manifolds, incorporating EM algorithm.
result PGPCA outperforms PPCA in modeling data around nonlinear manifolds.
New method for reducing dimensions of distributional data.
problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.
Dimensionality reduction is an important operation in information visualization, feature extraction, clustering, regression, and classification, especially for processing noisy high dimensional data. However, most existing approaches preserve either the global or the local structure of the data, but not both. Approache…
Proposes a deep learning method for effective data representation.
problem Constructing effective data representations for prediction.
method A deep dimension reduction approach to learning representations with sufficiency, low dimensionality, and disentanglement.
result The proposed deep nonparametric representation is consistent and performs better than existing methods.
Rdimtools simplifies DR and IDE for high-dimensional data analysis.
problem Discovering patterns in complex high-dimensional data.
method Provides an R package with 133 DR and 17 IDE algorithms.
result Facilitates geometric understanding of high-dimensional data.
RCLA reduces noise in topological data analysis, preserving essential structure.
problem Noise in large datasets obscures topological features in persistent homology.
method Grid-based RCLA integrates data reduction and denoising with a threshold parameter.
result RCLA provides a theoretical guarantee and automatic parameter selection.
SDR outperforms IDR in multimodal data analysis, especially with fewer samples.
problem Understanding and optimizing data efficiency in multimodal representation learning.
method Generative linear model to synthesize multimodal data, comparing IDR and SDR methods.
result Linear SDR methods yield higher-quality, more succinct reduced-dimensional representations with smaller datasets.
Proposes FMPCA for federated tensor data dimensionality reduction.
problem Integration of MPCA into federated learning.
method Federated Multilinear Principal Component Analysis (FMPCA).
result FMPCA preserves performance of traditional MPCA in federated learning.
Supervised dimensionality reduction strategies have been of great interest. However, current supervised dimensionality reduction approaches are difficult to scale for situations characterized by large datasets given the high computational complexities associated with such methods. While stochastic approximation strateg…
IKD uses eigen-decomposition for nonlinear dimensionality reduction.
problem Lack of sophisticated and nonlinear dimensionality reduction methods.
method Inverse Kernel Decomposition (IKD) based on eigen-decomposition of sample covariance matrix.
result IKD achieves comparable performance to optimization-based methods with faster running speeds.
A novel supervised visualization technique for data exploration.
problem Lack of supervised dimensionality reduction methods considering class labels.
method Random forest proximities and diffusion-based dimensionality reduction.
result Retains local and global structures in data, emphasizing important variables.
BlosSOM improves data visualization for large datasets.
problem Insufficient performance of dimensionality reduction methods for large datasets.
method GPU-accelerated semi-supervised EmbedSOM algorithm.
result Produces high-quality visualizations with user control.