This work refines Cover's theory for binary classification on low-dimensional data.
problem The challenge of analyzing how low-dimensional data structures affect classification models.
method Refines Cover's function-counting theory to account for low-dimensional data structure.
result Derives dichotomy counts and analyzes the impact of data structure on classification models.
New DR method uses Gromov-Wasserstein distance for high-dimensional data.
problem Analyzing relationships between high-dimensional objects.
method Optimal transportation theory and Gromov-Wasserstein distance.
result Robust and efficient solution for complex high-dimensional datasets.
PCENet reduces uncertainty in high-dimensional data efficiently.
problem Uncertainty quantification in high-dimensional data is computationally expensive.
method Two-stage learning process: variational autoencoder for low-dimensional representation, polynomial chaos expansion for mapping.
result Model captures system dynamics, learns under uncertainty, estimates high-dimensional data uncertainty, matches output distribution moments.
New method visualizes noisy data better than existing techniques.
problem Noisy data impairs data visualization methods.
method Functional Information Geometry (FIG) adapts EIG framework using functional data analysis.
result FIG outperforms EIG variant in capturing true structure, robustness, and speed.
GTSNE improves data visualization for high-dimensional data.
problem Visualizing high-dimensional data points in a 2D map.
method GTSNE is a variation of t-SNE that captures both local and macro structures.
result GTSNE produces better visualizations of high-dimensional data compared to other methods.
High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of dimensionality' phenomenon. Therefore, dimensionality reduction methods are applied to the…
Geometric framework detects outliers in high-dimensional data.
problem Detecting outliers in high-dimensional data.
method Geometric framework exploiting manifold structure.
result Significant improvement in outlier detection in high-dimensional data.
Anomaly detection for high-dimensional data using large deviations principle.
problem Challenges in anomaly detection for high-dimensional data.
method Large Deviations Anomaly Detection (LAD) algorithm.
result Outperforms state-of-the-art methods on high-dimensional data sets.
Generative models learn complex data from low-dimensional manifolds.
problem Theoretical justification for generative models on manifold structures.
method Prove statistical guarantees of generative networks under Wasserstein-1 loss, considering intrinsic dimensionality.
result Generative networks converge to zero at a fast rate depending on intrinsic dimensionality, not ambient data dimension.
Method determines latent dimensionality in international trade flows.
problem Finding meaningful low-dimensional latent features in high-dimensional international trade data.
method Proposes a latent dimension determination method based on clustering of nonnegative RESCAL decompositions.
result Validates the latent features against empirical economic facts.
GDMaps reduces high-dimensional data to lower dimensions for better classification.
problem High-dimensional data classification and representation.
method Grassmannian Diffusion Maps technique for nonlinear dimensionality reduction.
result GDMaps effectively identifies intrinsic subspace structures in high-dimensional data.
The paper finds non-Gaussian directions in high-dimensional data using Wasserstein distance.
problem Locating interesting non-Gaussian features in high-dimensional data.
method Projection pursuit using 2-Wasserstein distance to maximize the difference from Gaussian.
result Statistical guarantees for accurately approximating an unknown low-dimensional non-Gaussian subspace.
New method interpolates high-dimensional scattered data using kernel theory.
problem Scattered data in high-dimensional spaces defy traditional distributional assumptions.
method Kernel interpolation framework based on integral operator theory.
result Spectra of kernel matrices predict performance of interpolation methods.
A new method reduces high-dimensional data's impact on CWMs using TSNE.
problem High-dimensional data hampers CWMs' accuracy and speed.
method TSNE for dimensionality reduction, parsimonious technique, expectation maximization.
result TSNE enhances CWMs' performance in high-dimensional space.
Modern techniques simplify complex high-dimensional data.
problem Complex, high-dimensional data.
method Unsupervised dimension reduction techniques.
result Simplified representation of high-dimensional data.
Subspace clustering refers to the problem of clustering unlabeled high-dimensional data points into a union of low-dimensional linear subspaces, assumed unknown. In practice one may have access to dimensionality-reduced observations of the data only, resulting, e.g., from "undersampling" due to complexity and speed con…
New method estimates intrinsic dimensionality using angles, not distances.
problem Estimating local intrinsic dimensionality accurately.
method Introduces a new estimator using the distribution of angles between neighbor points.
result New estimator behaves similarly but complementarily to existing measures of intrinsic dimensionality.
This paper improves diffusion models for low-dimensional data.
problem Theoretical foundations of diffusion models are lacking for low-dimensional data.
method Score approximation, estimation, and distribution recovery of diffusion models on low-dimensional data.
result Sample complexity bounds for distribution estimation using diffusion models are provided.
Proposes generating virtual data points to overcome the curse of dimensionality.
problem Increased intrinsic dimensionality requires large data sets for local sampling.
method Manifold embedding motivated super sampling (MESS) framework.
result Generates virtual data points that faithfully represent the manifold.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
HD-BWDM improves clustering validation in high-dimensional data.
problem Determining the right number of clusters in high-dimensional data.
method HD-BWDM integrates random projection, PCA, trimmed clustering, and medoid-based distances.
result HD-BWDM remains stable and interpretable under high-dimensional projections and contamination.
Real world data often exhibit low-dimensional geometric structures, and can be viewed as samples near a low-dimensional manifold. This paper studies nonparametric regression of Hölder functions on low-dimensional manifolds using deep ReLU networks. Suppose n training data are sampled from a Hölder function in $\mathc…
Review of privacy-preserving linear models for high-dimensional data.
problem Overfitting and data memorization in high-dimensional linear models.
method Comprehensive comparison of optimization techniques for differentially private high-dimensional linear models.
result Coordinate-optimized algorithms perform best in empirical tests.
Proposes a deep neural network for multi-dimensional functional data classification.
problem Classifying multi-dimensional functional data with non-Gaussian distributions.
method Trains a deep neural network on the principle components of the training data.
result FDNN achieves minimax optimality when log density ratio has a locally connected modular structure.
This study evaluates clustering algorithms on high-dimensional data.
problem Comparing clustering algorithms on high-dimensional datasets.
method Evaluation of K-means, DBSCAN, and Spectral Clustering using PCA, t-SNE, UMAP, and multiple metrics.
result UMAP preprocessing improves clustering quality across all algorithms, with Spectral Clustering excelling.
Mercat preserves angles to create accurate low-dimensional embeddings.
problem Reconstructing global relationships in low-dimensional embeddings.
method Reconstructing angles between data points to preserve both local and global structures.
result Mercat yields good reconstruction across various experiments and metrics.
Kernel test detects manifold data differences with high-dimensional noise.
problem Detecting differences between manifold data samples.
method Kernel-based two-sample test statistic related to MMD for manifold data.
result The test power exceeds a threshold depending on manifold dimensionality, Hölder order, and squared divergence.
SSNL improves simulation-based inference for high-dimensional data.
problem Performance degradation in neural likelihood estimation for high-dimensional data.
method Surjective Sequential Neural Likelihood (SSNL) using surjective normalizing flow models.
result SSNL avoids manual crafting of summary statistics and outperforms state-of-the-art methods.
DeepFS uses deep neural networks to select significant features in ultra high-dimensional data.
problem Challenges in traditional feature selection methods for high-dimensional, low-sample-size data.
method Two-step nonparametric approach combining deep neural networks and feature screening.
result DeepFS effectively identifies significant features with high precision for ultra high-dimensional data.
Study improves Hayashi-Yoshida estimator for high-dimensional stock covolatility.
problem Inconsistent performance of Hayashi-Yoshida estimator in high dimensions.
method Analyzed the limiting spectral distribution of the Hayashi-Yoshida estimator.
result Established the connection between the estimator's spectrum and the true covariance matrix in high dimensions.
For low-dimensional data sets with a large amount of data points, standard kernel methods are usually not feasible for regression anymore. Besides simple linear models or involved heuristic deep learning models, grid-based discretizations of larger (kernel) model classes lead to algorithms, which naturally scale linear…
A new method identifies critical transitions in high-dimensional data.
problem Challenges in identifying critical transitions in high-dimensional time-series data.
method Spatial-temporal Principal Component Analysis (stPCA)
result Identifies tipping points before critical transitions reliably.
A new method improves graph-based learning for high-dimensional data.
problem Inconsistent high-dimensional learning efficiency of semi-supervised graph regularization.
method Introducing a novel regularization approach involving centering operation.
result Empirical results show improved performance over spectral clustering.
New algorithms extract low-dimensional representations from sequential data, revealing insights into complex processes.
problem Challenges in extracting low-dimensional representations from sequential, high-dimensional, sparse, and noisy data.
method Developed new clustering algorithms based on Block Markov Chains theory, validated on real-world data.
result These algorithms can successfully extract low-dimensional representations from real-world sequential data, revealing insights into complex processes.
Paper proposes data quality measures for large-scale high-dimensional data.
problem Lack of practical data quality measures for large-scale high-dimensional data.
method Proposes two data quality measures: class separability and in-class variability. Efficient algorithms based on random projections and bootstrapping are provided.
result Efficient algorithms for computing data quality measures on large-scale high-dimensional data.
Method learns low-dim. state vars from noisy high-dim. data.
problem Discovering dynamical models from noisy high-dimensional data.
method Stochastic Variational Deep Kernel Learning with encoder and latent model.
result Effective denoising, compact state representation, and uncertainty quantification.
Unified model for reducing dimensions and clustering high-dimensional data.
problem High-dimensional data clustering and dimensionality reduction.
method Hierarchical mixtures of Gaussians (HMoGs) with closed-form likelihood and inference.
result Efficiently models hundreds of latent dimensions, improving clustering performance.
Proposes a new method for causal inference in high-dimensional complex data.
problem Challenges in making causal inference with high-dimensional, nonlinear data.
method Combines deep learning techniques like sparse deep learning and stochastic neural networks.
result Outperforms existing methods in numerical studies.
Proposes GPLFR for predicting high-dimensional outputs with few data.
problem Predicting high-dimensional outputs from limited data.
method GPLFR combines Gaussian process and linear-Gaussian decoding for high-dimensional prediction.
result GPLFR outperforms existing methods in predicting high-dimensional outputs.
Scientific and engineering processes deliver massive high-dimensional data sets that are generated as non-linear transformations of an initial state and few process parameters. Mapping such data to a low-dimensional manifold facilitates better understanding of the underlying processes, and enables their optimization. I…
SDR outperforms IDR in multimodal data analysis, especially with fewer samples.
problem Understanding and optimizing data efficiency in multimodal representation learning.
method Generative linear model to synthesize multimodal data, comparing IDR and SDR methods.
result Linear SDR methods yield higher-quality, more succinct reduced-dimensional representations with smaller datasets.
For manifold learning, it is assumed that high-dimensional sample/data points are embedded on a low-dimensional manifold. Usually, distances among samples are computed to capture an underlying data structure. Here we propose a metric according to angular changes along a geodesic line, thereby reflecting the underlying …
Generative models can approximate high-dimensional data from lower dimensions without needing a latent dimension equal to or greater than the data's intrinsic dimension.
problem Theoretical limitations on the latent dimension required for generative models to approximate high-dimensional data distributions.
method Inspired by space-filling curves, the work demonstrates that generative networks can approximate distributions on d-dimensional manifolds from inputs of any arbitrary dimension, even lower than d. result Generative models can approximate high-dimensional data distributions from lower-dimensional inputs without needing a latent dimension equal to or greater than the data's intrinsic dimension.
The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…
rags2ridges simplifies graphical modeling of high-dimensional data.
problem Graphical modeling of high-dimensional precision matrices.
method Modular framework for extraction, visualization, and analysis of Gaussian graphical models.
result Provides a one-stop-shop for graphical modeling of high-dimensional precision matrices.
In order to avoid the curse of dimensionality, frequently encountered in Big Data analysis, there was a vast development in the field of linear and nonlinear dimension reduction techniques in recent years. These techniques (sometimes referred to as manifold learning) assume that the scattered input data is lying on a l…
Finding rare information hidden in a huge amount of data from the Internet is a necessary but complex issue. Many researchers have studied this issue and have found effective methods to detect anomaly data in low dimensional space. However, as the dimension increases, most of these existing methods perform poorly in de…
New method estimates treatment effects from high dimensional data.
problem Estimating treatment effects from high dimensional data with confounders.
method Generative modeling approach to backdoor adjustment in variational inference.
result Empirically, estimates interventional likelihood in high dimensional settings.