Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2905808691,159 · Jun 202019922001200920172026
48 results for Intrinsically low-dimensional data

New research shows DDPM can adapt to data's intrinsic low dimensionality efficiently.

problem Theoretical inefficiency of DDPM in high-dimensional data.
method Investigates how DDPM can exploit intrinsic low dimensionality of data.
result Proves DDPM's iteration complexity scales nearly linearly with intrinsic dimension kk.

Low-dimensional structure in images helps deep learning models generalize better.

problem Understanding the intrinsic dimensionality of images for better model performance.
method Applied dimension estimation tools to popular image datasets and used GANs to manipulate intrinsic dimensionality.
result Natural image datasets have very low intrinsic dimensionality, which aids neural networks in learning and generalizing.

Paper examines global Covid-19 data complexity and finds low intrinsic dimensions.

problem Understanding the complexity of Covid-19 data across countries.
method Used a Bayesian mixture model (Hidalgo) to estimate intrinsic dimensionality.
result Covid-19 data projects onto two low-dimensional manifolds without significant loss of information.

Wasserstein Autoencoders improve model efficiency and interpretability for low-dimensional data.

problem Limited statistical guarantees for WAEs in low-dimensional data.
method Proper network architecture selection and analysis of expected excess risk convergence rates.
result WAEs can learn data distributions efficiently when intrinsic dimension is considered.

Generative models learn complex data from low-dimensional manifolds.

problem Theoretical justification for generative models on manifold structures.
method Prove statistical guarantees of generative networks under Wasserstein-1 loss, considering intrinsic dimensionality.
result Generative networks converge to zero at a fast rate depending on intrinsic dimensionality, not ambient data dimension.

Diffusion models learn multi-modal distributions with optimal efficiency.

problem Learning high-dimensional distributions with low-dimensional multi-modal structures.
method Score-based diffusion models, focusing on subgaussian distributions within subspaces.
result Diffusion models require O~(εk2)\widetilde{O}(\varepsilon^{-k \vee 2}) samples for 1-Wasserstein ε\varepsilon error, improving over prior guarantees.

The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…

2016-06-28abs ↗pdf ↗

Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.

problem Learning high-dimensional representations from multi-modal data.
method Multi-modal contrastive learning with temperature optimization.
result Contrastive learning adapts to intrinsic dimensions of data, not specified dimensions.

The paper refutes the manifold hypothesis for image data and proposes the union of manifolds hypothesis.

problem The manifold hypothesis fails to capture the structure of image data.
method Empirical verification of the union of manifolds hypothesis on image datasets.
result Image data lies on a disconnected set with varying intrinsic dimensions.

Study shows how diffusion models learn on low-dimensional manifolds.

problem Learning efficiency of diffusion models on manifolds.
method Analyzes denoising score matching with random feature neural networks.
result Sample complexity scales linearly with intrinsic dimension, not ambient dimension.

ConvResNets approximate Besov functions and classify on low-dimensional manifolds.

problem Lack of statistical theories for deep learning on high-dimensional data.
method Exploits low-dimensional geometric structures of real-world data sets using ConvResNets.
result ConvResNets can approximate Besov functions and learn classifiers with optimal excess risk.

The study explains transformer scaling laws using statistical and approximation theories.

problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.

Paper explores how Rectified Flow adapts to low-dimensional data.

problem Improving sampling efficiency in low-dimensional data.
method Investigates Rectified Flow's adaptation to low-dimensional support and introduces a stochastic version.
result Shows improved sampling efficiency with O(k/ε)O(k/\varepsilon) complexity.

Novelty search in low-dimensional space improves sample efficiency in exploration tasks.

problem Efficient exploration in complex environments with sparse rewards.
method Combines model-based and model-free objectives to learn a low-dimensional representation. Uses intrinsic novelty rewards based on nearest neighbor distances in this space.
result Our approach achieves more sample-efficient exploration compared to strong baselines on various tasks.

Deep multi-task learning benefits from low intrinsic dimensionality, leading to better generalization.

problem Improving generalization in deep multi-task learning with high-dimensional models.
method Parametrizing multi-task networks in a low-dimensional space using random expansions and weight compression.
result First non-vacuous generalization bounds for deep multi-task networks are derived.

Chart autoencoders learn latent features preserving manifold topology and geometry, with robust denoising capabilities.

problem Learning low-dimensional latent features of high-dimensional data sampled near a manifold.
method Chart autoencoders encode data into latent features on charts, preserving manifold topology and geometry.
result Chart autoencoders achieve a squared generalization error of n2d+2log4nn^{-\frac{2}{d+2}}\log^4 n under proper network architectures.

Paper adapts DDPM to low-dimensional structures in image distributions.

problem Understanding and adapting to low-dimensional structures in image distributions.
method Developed a novel set of analysis tools to characterize algorithmic dynamics.
result First theoretical demonstration that DDPM can adapt to unknown low-dimensional structures.

Deep networks can adapt to intrinsic dimensionality beyond domain constraints.

problem Approximating functions on low-dimensional manifolds with high-dimensional data.
method Two-layer compositions with ReLU activation, using dimensionality reducing feature maps.
result Near optimal approximation rates depend on the complexity of the dimensionality reducing map, not the ambient dimension.

Isometry regularizer improves autoencoder performance on manifold learning.

problem Bad generalization in autoencoders, especially extrinsic and intrinsic issues.
method Introduces an isometry regularizer that encourages the decoder to be an isometry and the encoder to be its pseudo-inverse.
result Isometry regularizer leads to better generalization and useful low-dimensional data representations.

New theory shows deep networks adapt to data's intrinsic dimensionality even when data isn't on a low-dimensional manifold.

problem Existing theories on deep nonparametric regression assume data lie on a low-dimensional manifold, which is often not the case in real-world applications.
method Introduces effective Minkowski dimension to characterize the intrinsic dimension of data subsets and proves sample complexity depends on this new complexity notation.
result Deep neural networks can adapt to the effective Minkowski dimension of data, circumventing the curse of dimensionality for moderate sample sizes.

Scattering networks maximize separation on low-dimensional data.

problem Maximizing separation capacity on low-dimensional datasets.
method Characterize and bound separation capacity for feature extractors, then apply to scattering networks with specific criteria.
result Design criteria for scattering networks to maximize separation on low-dimensional data.

Data living on manifolds commonly appear in many applications. Often this results from an inherently latent low-dimensional system being observed through higher dimensional measurements. We show that under certain conditions, it is possible to construct an intrinsic and isometric data representation, which respects an …

2018-06-01abs ↗pdf ↗

Paper improves deep learning convergence rates for low-dimensional data.

problem Sub-optimal rates in deep learning due to unrealistic assumptions on intrinsic dimension.
method Introduced an entropic notion of intrinsic dimension for exponential families and demonstrated improved convergence rates.
result Test error scales as O~(n2β2β+dˉ2β(λ))\tilde{\mathcal{O}}\left(n^{-\frac{2β}{2β+ \bar{d}_{2β}(λ)}}\right), improving on best-known rates.

Paper analyzes dataset distillation for efficient encoding of task-relevant information.

problem Efficiently encoding task-relevant information from gradient-based learning of non-linear tasks.
method Theoretical analysis of dataset distillation applied to two-layer neural networks with gradient-based training.
result Low-dimensional structure of the problem is efficiently encoded into distilled data, reproducing a model with high generalization ability.

The paper improves the probability flow ODE sampler for faster sampling of natural images.

problem Improving the convergence rate of the probability flow ODE sampler.
method Adapting the probability flow ODE sampler to exploit intrinsic low-dimensional structures in natural image data.
result Achieves a dimension-free convergence rate of O(k/T)O(k/T) in total variation distance, improving upon existing results.

This paper describes a method for learning low-dimensional approximations of nonlinear dynamical systems, based on neural-network approximations of the underlying Koopman operator. Extended Dynamic Mode Decomposition (EDMD) provides a useful data-driven approximation of the Koopman operator for analyzing dynamical syst…

2017-12-04abs ↗pdf ↗

Recent theory work has found that a special type of spatial partition tree - called a random projection tree - is adaptive to the intrinsic dimension of the data from which it is built. Here we examine this same question, with a combination of theory and experiments, for a broader class of trees that includes k-d trees…

2012-05-09abs ↗pdf ↗

This paper analyzes deep federated learning for low-dimensional data, revealing intrinsic dimensionality's role in convergence rates.

problem Insufficient investigation of generalization error in heterogeneous federated learning, especially for low-dimensional data.
method Statistical analysis of deep federated regression in a two-stage sampling model.
result Intrinsic dimensionality, characterized by entropic dimension, determines convergence rates for deep learners.

New diffusion models learn distributions from samples with improved error bounds.

problem Statistical guarantees for score-based diffusion models on low-dimensional data.
method Derive finite-sample error bounds for Wasserstein-pp distance.
result Error bounds scale as n1/dp,q(μ)n^{-1 / d^\ast_{p,q}(μ)} for diffusion models.

A new method for SVGD reduces variance in high dimensions.

problem High-dimensional variance in SVGD.
method Grassmann Stein Variational Gradient Descent (GSVGD) projects onto arbitrary subspaces and uses coupled Grassmann-valued diffusion.
result GSVGD explores high-dimensional problems with intrinsic low-dimensional structure efficiently.

In this paper, we build an organization of high-dimensional datasets that cannot be cleanly embedded into a low-dimensional representation due to missing entries and a subset of the features being irrelevant to modeling functions of interest. Our algorithm begins by defining coarse neighborhoods of the points and defin…

2015-07-01abs ↗pdf ↗

The paper improves GANs' theoretical guarantees for low-dimensional data.

problem Theoretical guarantees for GANs' statistical accuracy remain pessimistic.
method Analytical derivation of statistical guarantees on estimated densities.
result Theoretical rates of convergence for GANs and BiGANs are derived.

TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.

problem Stochastic linear bandits with partially observed actions in settings like recommendation and healthcare.
method TOFU-POV estimates latent action subspace, imputes missing actions, and runs OFUL in low-dimensional coordinates.
result TOFU-POV achieves T\sqrt{T} regret scaling with intrinsic subspace dimension, improving upon natural baselines.

The paper analyzes reflected diffusion models on hypercube data.

problem Challenges in modeling bounded domains with low-dimensional data.
method Employed an infinite series expansion of transition densities to bound the score function and its approximation.
result Established convergence rates for generative algorithm adapting to intrinsic dimensionality.

A new method integrates autoencoders with geometry regularization for manifold learning.

problem Extracting simplified low-dimensional representations that capture intrinsic geometry in data.
method Integrates autoencoders with a geometric regularization term based on diffusion potential distances.
result The method preserves intrinsic structure, enables out-of-sample extension, and faithful reconstruction.

Deep networks can approximate high-dimensional distributions from low-dimensional ones.

problem Approximating high-dimensional distributions from low-dimensional ones.
method Proved neural networks can transform low-dimensional distributions to high-dimensional ones with arbitrary closeness measured by Wasserstein distances and maximum mean discrepancy.
result Upper bounds of the approximation error are obtained in terms of the width and depth of neural network.

W2S FT often outperforms weak teachers due to low intrinsic dimensionality.

problem Understanding why weak-to-strong finetuning outperforms weak models.
method Analyzing W2S in ridgeless regression setting, focusing on variance reduction.
result Weak teacher's variance is inherited by strong student in shared feature subspace, reduced in discrepancy subspace.

Kernel-spectral embedding learns low-dim. structures from noisy data.

problem Learning low-dimensional nonlinear structures from high-dimensional noisy data.
method Adaptive bandwidth spectral embedding using integral operators.
result Convergence to noiseless embeddings and eigenfunctions of integral operators.