Efficient method estimates intrinsic dimension for big data.
problem Estimating intrinsic dimension for large datasets is costly and complex.
method Proposes a matrix-vector product-based approach for efficient intrinsic dimension estimation.
result Demonstrates superior performance compared to state-of-the-art methods.
Rdimtools simplifies DR and IDE for high-dimensional data analysis.
problem Discovering patterns in complex high-dimensional data.
method Provides an R package with 133 DR and 17 IDE algorithms.
result Facilitates geometric understanding of high-dimensional data.
eDCF estimates intrinsic dimension using local connectivity.
problem Challenges in estimating intrinsic dimension due to scale dependence.
method eDCF: a novel, scalable, and parallelizable method based on Connectivity Factor (CF).
result eDCF consistently matches leading estimators with comparable MAE and higher exact intrinsic dimension match rates.
Data-driven method solves multiscale elliptic PDEs with random coefficients.
problem Solving multiscale elliptic PDEs with random coefficients.
method Data-driven approach based on intrinsic dimension reduction.
result Efficient solution of multiscale elliptic PDEs with random coefficients.
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
problem Efficiency in handling large datasets with high intrinsic dimensions.
method CCP partitions features into correlated clusters and projects them to 1D based on sample correlations.
result CCP achieves efficient dimensionality reduction without matrix diagonalization.
New algorithm estimates intrinsic dimension of discrete datasets.
problem Inaccuracies in using continuous methods for discrete datasets.
method Introduced an algorithm to infer intrinsic dimension of discrete spaces.
result Demonstrated accuracy on benchmark datasets and found a small intrinsic dimension in a metagenomic dataset.
LIDL estimates local intrinsic dimension in high dimensions.
problem Estimating local intrinsic dimension in high-dimensional data.
method Approximate likelihood using parametric neural density estimation.
result LIDL scales to thousands of dimensions and yields competitive results.
Sliced inverse regression (SIR) is a pioneer tool for supervised dimension reduction. It identifies the effective dimension reduction space, the subspace of significant factors with intrinsic lower dimensionality. In this paper, we propose to refine the SIR algorithm through an overlapping slicing scheme. The new algor…
This paper tackles high-dimensional Bayesian optimization using supervised dimension reduction.
problem Challenges in extending Bayesian optimization to high dimensions.
method Introduces Sliced Inverse Regression (SIR) for high-dimensional Bayesian optimization.
result Demonstrates computational benefits and theoretical regret bounds for high-dimensional Bayesian optimization.
Adaptive framework improves nonparametric dimensionality reduction.
problem Optimal hyper-parameter tuning for nonparametric dimensionality reduction.
method Adaptive framework using intrinsic dimension estimator and optimal local neighbourhood sizes.
result Significant improvements in various learning tasks through better low-dimensional visualizations.
This paper explores autoencoders for estimating intrinsic dimensionality.
problem Estimating the intrinsic dimensionality of random vectors.
method Use of autoencoders for dimension estimation, focusing on architectural choices and regularization techniques.
result Autoencoders can be adapted for intrinsic dimension estimation, addressing questions beyond classic DR/DE techniques.
We investigate quaternionic contact (qc) manifolds from the point of view of intrinsic torsion. We argue that the natural structure group for this geometry is a non-compact Lie group K containing Sp(n)H^*, and show that any qc structure gives rise to a canonical K-structure with constant intrinsic torsion, except in se…
We study adaptive data-dependent dimensionality reduction in the context of supervised learning in general metric spaces. Our main statistical contribution is a generalization bound for Lipschitz functions in metric spaces that are doubling, or nearly doubling. On the algorithmic front, we describe an analogue of PCA f…
Introduces intrinsic Riemannian cross-covariance for manifold-valued random objects.
problem Covariance estimation for random objects on Riemannian manifolds.
method Defines covariance and correlation via parallel transport.
result Proposed covariance is independent of coordinate choices.
In this expository paper we present short simple proofs of Conway-Gordon-Sachs' theorem on intrinsic linking in three-dimensional space, as well as van Kampen-Flores' and Ummel's theorems on intrinsic intersections. The latter are related to nonrealizability of certain hypergraphs in four-dimensional space. The proofs …
Information about intrinsic dimension is crucial to perform dimensionality reduction, compress information, design efficient algorithms, and do statistical adaptation. In this paper we propose an estimator for the intrinsic dimension of a data set. The estimator is based on binary neighbourhood information about the ob…
A new geometry-preserving method for interpreting compositional data.
problem Statistical challenges in high-dimensional compositional data.
method Geometry-preserving framework for dimension reduction of compositional data.
result Identification of a central compositional subspace for compositional predictors.
Deep networks can adapt to intrinsic dimensionality beyond domain constraints.
problem Approximating functions on low-dimensional manifolds with high-dimensional data.
method Two-layer compositions with ReLU activation, using dimensionality reducing feature maps.
result Near optimal approximation rates depend on the complexity of the dimensionality reducing map, not the ambient dimension.
Dimensionality reduction is a topic of recent interest. In this paper, we present the classification constrained dimensionality reduction (CCDR) algorithm to account for label information. The algorithm can account for multiple classes as well as the semi-supervised setting. We present an out-of-sample expressions for …
Coadjoint orbits for the group SO(6) parametrize Riemannian G-reductions in six dimensions, and we use this correspondence to interpret symplectic fibrations between these orbits, and to analyse moment polytopes associated to the standard Hamiltonian torus action on the coadjoint orbits. The theory is then applied to d…
Investigates projections onto explicit subspaces and their variance effects.
problem Understanding the variance preservation in explicit subspace projections.
method Investigates projections onto explicit subspaces of varying dimensionality and analyzes the variance effects.
result Developed new bounds for Euclidean distances and inner products.
This paper deals with a new filter algorithm for selecting the smallest subset of features carrying all the information content of a data set (i.e. for removing redundant features). It is an advanced version of the fractal dimension reduction technique, and it relies on the recently introduced Morisita estimator of Int…
W2S FT often outperforms weak teachers due to low intrinsic dimensionality.
problem Understanding why weak-to-strong finetuning outperforms weak models.
method Analyzing W2S in ridgeless regression setting, focusing on variance reduction.
result Weak teacher's variance is inherited by strong student in shared feature subspace, reduced in discrepancy subspace.
This research examines how data transformations affect adversarial robustness in recurrent neural networks.
problem Adversarial examples reduce machine learning accuracy, especially in high-dimensional datasets.
method Analysis of feature selection, dimensionality reduction, and trend extraction techniques on recurrent neural networks.
result Data transformations may increase vulnerability to adversarial samples, but only if they approximate intrinsic dimensionality and maintain manifold coverage.
We give an intrinsic definition of (affine very) special real manifolds and realise any such manifold M as a domain in affine space equipped with a metric which is the Hessian of a cubic polynomial. We prove that the tangent bundle N=TM carries a canonical structure of (affine) special Kähler manifold. This gives a…
The paper evaluates and compares dimensionality reduction quality metrics without tuning.
problem Evaluating the quality of nonlinear dimensionality reduction visualizations is challenging.
method Comparison of dimensionality reduction quality metrics on datasets with known ground truth manifolds.
result A few methods consistently perform well, with one proposed as a benchmark.
Estimates intrinsic dimension of data for GANs.
problem Estimating intrinsic dimension of high-dimensional data.
method Uses Wasserstein distances for estimation.
result Provides sample complexity bounds for GANs.
Dimensionality-reduction methods are a fundamental tool in the analysis of large data sets. These algorithms work on the assumption that the "intrinsic dimension" of the data is generally much smaller than the ambient dimension in which it is collected. Alongside their usual purpose of mapping data into a smaller dimen…
Survey on geometric foundations of data reduction methods.
problem High-dimensional data with intrinsic nonlinear structure.
method Spectral manifold learning methods.
result Derivation and convergence analysis of spectral manifold learning.
A new method for SVGD reduces variance in high dimensions.
problem High-dimensional variance in SVGD.
method Grassmann Stein Variational Gradient Descent (GSVGD) projects onto arbitrary subspaces and uses coupled Grassmann-valued diffusion.
result GSVGD explores high-dimensional problems with intrinsic low-dimensional structure efficiently.
Paper infers intrinsic dimension from quasi-convex measurements.
problem Inferring intrinsic dimension from measurements by quasi-convex functions.
method Developed a method using filtration of Dowker complexes based on discrete data of point orderings.
result Correct intrinsic dimension can be inferred in the limit of large data under generic assumptions.
New estimators for intrinsic dimension and Wasserstein distance improve OT accuracy.
problem Intrinsic dimension estimation and Wasserstein distance estimation in large-scale OT.
method Introduces novel estimators for intrinsic dimension and Wasserstein distance.
result Simple, tuning-free estimator of OT and fast intrinsic dimension estimator.
Python package for estimating intrinsic dimensionality of datasets.
problem Estimating intrinsic dimensionality for machine learning applications.
method Implementation of various intrinsic dimension estimators in Python.
result Benchmarking of ID estimation methods on real and synthetic data.
Compressing data helps learn Mahalanobis metrics effectively.
problem Learning Mahalanobis metrics in high-dimensional spaces.
method Randomly compress data to train a full-rank metric in a reduced feature space.
result Theoretical guarantees on error for Mahalanobis metric learning, independent of ambient dimension.
GDMaps reduces high-dimensional data to lower dimensions for better classification.
problem High-dimensional data classification and representation.
method Grassmannian Diffusion Maps technique for nonlinear dimensionality reduction.
result GDMaps effectively identifies intrinsic subspace structures in high-dimensional data.
Many nonparametric regressors were recently shown to converge at rates that depend only on the intrinsic dimension of data. These regressors thus escape the curse of dimension when high-dimensional data has low intrinsic dimension (e.g. a manifold). We show that k-NN regression is also adaptive to intrinsic dimension. …
The paper corrects biases in estimating intrinsic dimension and differential entropy.
problem Systematic bias in estimating intrinsic dimension and differential entropy.
method A bias-corrected estimator for both measures is proposed, highlighting shared steps and useful consequences.
result Simultaneous estimation of differential entropy and intrinsic dimension provides complementary perspectives on underlying manifolds.
Deep networks learn low-dimensional yet complex data representations.
problem Understanding the intrinsic dimensionality of deep neural network representations.
method Analysis of intrinsic dimensionality across multiple layers of trained networks.
result The intrinsic dimensionality of data representations in deep networks is significantly lower than the number of units in each layer.
Study provides bounds for estimating intrinsic dimension using Gaussian kernels.
problem Estimating intrinsic dimension from data.
method Finite-sample concentration and anti-concentration bounds for Gaussian kernel sums.
result Explicit dependence on sample size, bandwidth, and geometric parameters.
Rollings of reductive homogeneous spaces are studied using intrinsic curves.
problem Investigate rollings of reductive homogeneous spaces without slip and twist.
method An intrinsic point of view, considering rollings as curves in the configuration space Q tangent to a certain distribution. result Explicit solutions for rollings of m over G/H are obtained for specific cases. Unified analysis simplifies Johnson-Lindenstrauss lemma for data reduction.
problem Efficiently reducing high-dimensional data while preserving geometry.
method Unified analysis of various JL constructions using probabilistic tools.
result First rigorous proof and extension of spherical construction's effectiveness.
We analyse the relationship between the components of the intrinsic torsion of an SU(3) structure on a 6-manifold and a G_2 structure on a 7-manifold. Various examples illustrate the type of SU(3) structure that can arise as a reduction of a metric with holonomy G_2.
Study on rolling Stiefel manifolds with specific metrics.
problem Intrinsic and extrinsic rolling of Stiefel manifolds with α-metrics. method Investigation of intrinsic rolling of normal naturally reductive homogeneous spaces, derivation of ODEs for rolling, and explicit solutions.
result Explicit solutions for intrinsic and extrinsic rolling of Stiefel manifolds.
New research shows that the dimension gap between intrinsic and ambient dimensions affects adversarial vulnerability of machine learning models.
problem The mystery of adversarial attacks on machine learning models.
method Introducing two types of adversarial attacks and proving their relationship to the dimension gap.
result The dimension gap between intrinsic and ambient dimensions makes clean-trained models more vulnerable to off-manifold adversarial perturbations.
Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.
problem Learning high-dimensional representations from multi-modal data.
method Multi-modal contrastive learning with temperature optimization.
result Contrastive learning adapts to intrinsic dimensions of data, not specified dimensions.
We propose a new method for estimating the intrinsic dimension of a dataset by applying the principle of regularized maximum likelihood to the distances between close neighbors. We propose a regularization scheme which is motivated by divergence minimization principles. We derive the estimator by a Poisson process appr…
Study shows DNNs perform well with low intrinsic data dimensions.
problem Understanding DNN performance with high-dimensional data.
method Derived bounds for approximation and generalization errors, developed novel proof technique.
result Convergence rates of DNN errors are independent of high dimensionality but dependent on intrinsic low dimensionality.
Low-dimensional structure in images helps deep learning models generalize better.
problem Understanding the intrinsic dimensionality of images for better model performance.
method Applied dimension estimation tools to popular image datasets and used GANs to manipulate intrinsic dimensionality.
result Natural image datasets have very low intrinsic dimensionality, which aids neural networks in learning and generalizing.