This work refines Cover's theory for binary classification on low-dimensional data.
problem The challenge of analyzing how low-dimensional data structures affect classification models.
method Refines Cover's function-counting theory to account for low-dimensional data structure.
result Derives dichotomy counts and analyzes the impact of data structure on classification models.
Deep ReLU networks estimate Hölder functions on low-dimensional manifolds with fast convergence.
problem Estimating Hölder functions on low-dimensional manifolds with noisy data.
method Deep ReLU network architecture designed for nonparametric regression.
result Empirical estimator convergence rate of n − 2 ( s + α ) 2 ( s + α ) + d log 3 n n^{-\frac{2(s+α)}{2(s+α) + d}}\log^3 n n − 2 ( s + α ) + d 2 ( s + α ) log 3 n . Generative models learn complex data from low-dimensional manifolds.
problem Theoretical justification for generative models on manifold structures.
method Prove statistical guarantees of generative networks under Wasserstein-1 loss, considering intrinsic dimensionality.
result Generative networks converge to zero at a fast rate depending on intrinsic dimensionality, not ambient data dimension.
Diffusion models adapt to low-dimensional data regardless of coefficient choices.
problem Understanding how diffusion models adapt to low-dimensional data structures.
method Analysis of diffusion models with flexible coefficient choices.
result Proven that O ~ ( k / ε ) \widetilde{O}(k/\varepsilon) O ( k / ε ) iterations suffice for accurate sampling in total variation distance. The paper analyzes side effects of learning from low-dimensional data embedded in a Euclidean space.
problem Learning from data distributed in a linear subspace of high-dimensional space.
method Derives estimates on the variation of the learning function and studies regularization effects.
result Potential regularization effects associated with network depth and noise in codimension of data manifold.
This paper improves diffusion models for low-dimensional data.
problem Theoretical foundations of diffusion models are lacking for low-dimensional data.
method Score approximation, estimation, and distribution recovery of diffusion models on low-dimensional data.
result Sample complexity bounds for distribution estimation using diffusion models are provided.
Paper solves globally optimal k-means for low dimensional data.
problem Finding globally optimal k-means solutions for low dimensional data.
method Formulates as a concave assignment problem, iteratively solving small concave and large linear programming problems.
result Solves k-means to global optimality for large data sets with several clusters.
Paper uses low-dimensional sensor data analysis for better fault detection.
problem Fault detection in critical equipment using multivariate, nonlinear sensor data.
method Exploits t-SNE and KPCA for nonlinear dimension reduction and anomaly detection.
result Low-dimensional representations improve interpretability and edge processing in IoT.
New algorithms extract low-dimensional representations from sequential data, revealing insights into complex processes.
problem Challenges in extracting low-dimensional representations from sequential, high-dimensional, sparse, and noisy data.
method Developed new clustering algorithms based on Block Markov Chains theory, validated on real-world data.
result These algorithms can successfully extract low-dimensional representations from real-world sequential data, revealing insights into complex processes.
ConvResNets approximate Besov functions and classify on low-dimensional manifolds.
problem Lack of statistical theories for deep learning on high-dimensional data.
method Exploits low-dimensional geometric structures of real-world data sets using ConvResNets.
result ConvResNets can approximate Besov functions and learn classifiers with optimal excess risk.
The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…
We consider the ability of deep neural networks to represent data that lies near a low-dimensional manifold in a high-dimensional space. We show that deep networks can efficiently extract the intrinsic, low-dimensional coordinates of such data. We first show that the first two layers of a deep network can exactly embed…
New research shows DDPM can adapt to data's intrinsic low dimensionality efficiently.
problem Theoretical inefficiency of DDPM in high-dimensional data.
method Investigates how DDPM can exploit intrinsic low dimensionality of data.
result Proves DDPM's iteration complexity scales nearly linearly with intrinsic dimension k k k . Two-layer networks trained on low-dimensional subspaces are vulnerable to adversarial examples.
problem Vulnerability of two-layer neural networks to adversarial examples on low-dimensional subspaces.
method Analysis of gradient behavior and effect of initialization scale and regularization.
result Decreasing initialization scale or adding L2 regularization can improve robustness to adversarial perturbations orthogonal to the data.
Entropy-Isomap improves low-dimensional visualization of dynamic processes.
problem Mapping high-dimensional, temporally correlated data to a low-dimensional manifold.
method Entropy-Isomap, a novel method addressing temporal correlations in data.
result Correctly captures process control variables and material morphology evolution.
The paper develops a neural network for learning dynamics from data.
problem Trade-off between representational capacity and overfitting in EDMD.
method Linear recurrent autoencoder network for Koopman operator approximation.
result Improved model reduction and nonlinear reconstruction techniques.
UAPCA projects uncertain data to low dimensions using GMMs.
problem Uncertain multidimensional data not well described by normal distributions.
method Model data with Gaussian mixture models, derive UAPCA projection from general formulation.
result Low-dimensional projections better represent multidimensional distributions.
Cryo-em images are found to be low-dimensional.
problem Understanding the geometric structure of cryo-em data.
method Applied manifold learning techniques to CryoSBI representations.
result Cryo-em data inherently populate low-dimensional manifolds.
New method identifies differences between groups in low-dimensional data representations.
problem Identifying meaningful differences between groups in low-dimensional data representations.
method Introduce Global Counterfactual Explanation (GCE) and Transitive Global Translations (TGT) for computing GCEs.
result TGT identifies sparse, accurate explanations that match real data patterns.
Paper provides statistical guarantees for GANs estimating Hölder space densities.
problem Statistical properties and theoretical guarantees for GANs.
method Approximation and statistical guarantees for GANs using Hölder space densities.
result GANs are consistent estimators of data distributions under strong discrepancy metrics.
Extends dimension reduction to data-driven settings without gradients.
problem Gradient-based dimension reduction limitations in data-driven settings.
method Score ratio matching framework, tailored parameterization, regularization, eigenvalue deflation.
result Outperforms standard score-matching for problems with low-dimensional structure.
Scattering networks maximize separation on low-dimensional data.
problem Maximizing separation capacity on low-dimensional datasets.
method Characterize and bound separation capacity for feature extractors, then apply to scattering networks with specific criteria.
result Design criteria for scattering networks to maximize separation on low-dimensional data.
Study shows how diffusion models learn on low-dimensional manifolds.
problem Learning efficiency of diffusion models on manifolds.
method Analyzes denoising score matching with random feature neural networks.
result Sample complexity scales linearly with intrinsic dimension, not ambient dimension.
A method for low-dimensional MDP representation using deep neural networks.
problem Constructing a low-dimensional Markov decision process representation.
method Use a deep neural network to define a class of potential process representations and estimate the process of lowest dimension.
result A decision strategy that maximizes mean utility for the low-dimensional representation also maximizes mean utility for the original process.
Two methods factor out prior knowledge from low-dimensional embeddings.
problem Visualizing data without considering background knowledge.
method JEDI for tSNE and CONFETTI for any embedding.
result Embeddings reveal meaningful structure hidden by prior knowledge.
Method identifies key features for clustering in high-dimensional data.
problem Understanding hidden patterns in high-dimensional data.
method Unsupervised feature selection based on discriminative power.
result 27 key transcription factors identified, 18 known to define cell states.
Improved likelihood estimation for singular distributions using deep models.
problem Estimating singular distributions using deep generative models.
method Data perturbation to avoid singularity issues in likelihood estimation.
result Consistent estimation of target distribution with desirable rates.
Paper explores how Rectified Flow adapts to low-dimensional data.
problem Improving sampling efficiency in low-dimensional data.
method Investigates Rectified Flow's adaptation to low-dimensional support and introduces a stochastic version.
result Shows improved sampling efficiency with O ( k / ε ) O(k/\varepsilon) O ( k / ε ) complexity. Mercat preserves angles to create accurate low-dimensional embeddings.
problem Reconstructing global relationships in low-dimensional embeddings.
method Reconstructing angles between data points to preserve both local and global structures.
result Mercat yields good reconstruction across various experiments and metrics.
We develop a method to summarize causal models with cycles in cubic time.
problem Cycles in high-dimensional causal models limit applicability of existing methods.
method We relax the acyclicity assumption in LiNG models and develop a low-dimensional DAG summary.
result Our method allows recovery of a low-dimensional DAG from high-dimensional data with cycles.
This paper improves conditional multidimensional scaling for incomplete data.
problem Handling missing data in known features for multidimensional scaling.
method Proposes a method to learn low-dimensional configurations with missing known feature values.
result Can learn low-dimensional configurations and impute missing values.
Kernel models learn low-dimensional predictive subspaces from input data.
problem Learning effective feature transformations in kernel models.
method Study of a compositional kernel ridge regression model.
result Global minimizers of the objective function identify the subspace with high probability.
Low-dimensional structure in images helps deep learning models generalize better.
problem Understanding the intrinsic dimensionality of images for better model performance.
method Applied dimension estimation tools to popular image datasets and used GANs to manipulate intrinsic dimensionality.
result Natural image datasets have very low intrinsic dimensionality, which aids neural networks in learning and generalizing.
A new method for embedding data in low dimensions, robust to noise.
problem Finding high-quality embeddings in noisy data.
method Formulates embedding as a robust ranking problem over triplets.
result Produces better embeddings with less noise and faster computation.
Wasserstein Autoencoders improve model efficiency and interpretability for low-dimensional data.
problem Limited statistical guarantees for WAEs in low-dimensional data.
method Proper network architecture selection and analysis of expected excess risk convergence rates.
result WAEs can learn data distributions efficiently when intrinsic dimension is considered.
Extract low-dimensional dynamics from multiple neural recordings.
problem Current methods can't handle dynamics across multiple neural recordings.
method Subspace-identification approach with moment-matching objective and scalable stochastic gradient descent.
result Can identify dynamics and predict correlations even with missing data and small overlap.
Optimizes optimal transport distances using low-dimensional embeddings.
problem High computational cost of optimal transport distances in high dimensions.
method Approximate OT distances using 1-Lipschitz maps in a lower-dimensional space.
result Efficiently approximates optimal transport distances with lower computational cost.
Finding rare information hidden in a huge amount of data from the Internet is a necessary but complex issue. Many researchers have studied this issue and have found effective methods to detect anomaly data in low dimensional space. However, as the dimension increases, most of these existing methods perform poorly in de…
The paper explores how regularization can improve multi-objective learning with high-dimensional data.
problem Improving multi-objective learning with high-dimensional and costly data.
method A two-stage MOL framework that leverages low-dimensional structure.
result Vanilla regularization approaches often fail in multi-objective learning, and a two-stage framework can successfully exploit low-dimensional structure.
A new framework detects anomalies in structured data.
problem Detecting anomalies in samples not conforming to low-dimensional manifolds.
method Preference Isolation Forest (PIF) framework combining adaptive isolation methods and preference embedding.
result Anomalies identified as isolated points in a high-dimensional preference space.
Consider a dataset of vector-valued observations that consists of noisy inliers, which are explained well by a low-dimensional subspace, along with some number of outliers. This work describes a convex optimization problem, called REAPER, that can reliably fit a low-dimensional model to this type of data. This approach…
Kernel-spectral embedding learns low-dim. structures from noisy data.
problem Learning low-dimensional nonlinear structures from high-dimensional noisy data.
method Adaptive bandwidth spectral embedding using integral operators.
result Convergence to noiseless embeddings and eigenfunctions of integral operators.
Adaptive GMRA approximates high-dimensional data with low-dimensional geometric structures.
problem Efficiently approximating high-dimensional data from a nearly low-dimensional manifold.
method Adaptive GMRA with thresholding of geometric wavelet coefficients.
result Adaptive GMRA approximations perform well on various measures with different regularity.
Paper adapts DDPM to low-dimensional structures in image distributions.
problem Understanding and adapting to low-dimensional structures in image distributions.
method Developed a novel set of analysis tools to characterize algorithmic dynamics.
result First theoretical demonstration that DDPM can adapt to unknown low-dimensional structures.
Study on low-dimensional adversarial perturbations in classification models.
problem Understanding and quantifying the effectiveness of low-dimensional adversarial perturbations.
method Analytical lower-bounds for fooling rate, considering binary classifiers under generic regularity conditions.
result Rigorous explanation for the success of heuristic methods in generating low-dimensional adversarial perturbations.
The study explains transformer scaling laws using statistical and approximation theories.
problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.
New method detects low-dimensional manifolds within bounds.
problem Detecting low-dimensional manifolds within bounds.
method Matrix Completion (MC) problem for partially observed distances on a manifold.
result The method provides theoretical guarantees on manifold detection and robustness to non-uniform sampling.
Paper explores tradeoff between standard and robust accuracy for latent models.
problem Tradeoff between standard accuracy and robust accuracy in adversarial training.
method Revisits adversarial training for latent models, considering Gaussian mixture and generalized linear models.
result Low-dimensional manifold structure mitigates the tradeoff between standard and robust accuracy.