New metrics for high-dimensional data improve on energy distance.
problem Testing equality of distributions and independence in high dimensions.
method Proposed new metrics inheriting properties of energy distance and others.
result Improved metrics detect homogeneity and independence in high dimensions.
Paper proposes new density estimators for high-dimensional data.
problem Prohibitive computational cost and slow convergence rate in high-dimensional density estimation.
method Adaptive hyperbolic cross density estimators in mixed smooth Sobolev spaces.
result Proposed estimators do not suffer curse of dimensionality under Integral Probability Metrics.
DMLreg uses expert knowledge to improve model performance in high-dimensional settings.
problem Improving model performance in high-dimensional prediction problems.
method Learning a Mahalanobis distance metric from expert comparisons and integrating it into a regularized linear model.
result DMLreg leads to improvements in model performance when expert knowledge is relevant.
Riemannian metric matching learns the geometry of high-dimensional datasets using neural networks.
problem Estimating the geometry of high-dimensional datasets from samples
method Riemannian metric matching using neural networks
result Riemannian metric matching rivals or improves k-NN-based diffusion geometry estimators New proof confirms nearly Sasakian condition for high-dimensional manifolds.
problem Characterizing Sasakian manifolds among almost contact metric manifolds.
method Self-contained, conceptual proof for nearly Sasakian condition.
result An almost contact metric manifold is Sasakian if and only if it is nearly Sasakian.
A new geometric metric identifies true data changes from parametrization artifacts in high-dimensional representations.
problem Quantifying representation drift in high-dimensional data using Euclidean or cosine distances can misattribute changes due to arbitrary parametrizations.
method Introducing the Fubini Study metric to identify representations that differ only by gauge transformations.
result The Fubini Study metric isolates intrinsic evolution by remaining invariant under gauge-induced fluctuations, providing a diagnostic for meaningful structural changes.
FLRML efficiently handles large-scale, high-dimensional data for metric learning.
problem Metric learning for large-scale, high-dimensional datasets is computationally expensive and memory-intensive.
method FLRML reformulates low-rank metric learning as an unconstrained optimization problem on the Stiefel manifold, enabling efficient mini-batch processing.
result FLRML achieves high accuracy with significantly reduced computational and memory costs.
In this article the package High-dimensional Metrics (\texttt{hdm}) is introduced. It is a collection of statistical methods for estimation and quantification of uncertainty in high-dimensional approximately sparse models. It focuses on providing confidence intervals and significance testing for (possibly many) low-dim…
This study evaluates clustering algorithms on high-dimensional data.
problem Comparing clustering algorithms on high-dimensional datasets.
method Evaluation of K-means, DBSCAN, and Spectral Clustering using PCA, t-SNE, UMAP, and multiple metrics.
result UMAP preprocessing improves clustering quality across all algorithms, with Spectral Clustering excelling.
Defines cross product for m vectors in n-dimensional spaces.
problem No universal definition for cross product in high-dimensional spaces.
method Defines cross product for m vectors in n-dimensional spaces with any metric matrices.
result Cross product length represents m-dimensional volume, components represent volume directions.
Deep metric learning detects anomalies without labels.
problem Unsupervised anomaly detection for high-dimensional data.
method Deep metric learning with end-to-end optimization, data distillation, hard mining.
result Significant performance gains over state-of-the-art methods.
Paper tackles blind detection of molecular signals in non-coherent media.
problem Long-tail channel response causes ISI, deteriorating detection performance.
method Develops a high-dimensional non-coherent scheme combining different non-coherent metrics.
result Higher dimensionality in metric space achieves lower bit error rate (BER).
Study high-dimensional covariance matrix estimators for complex portfolios, improving financial metrics.
problem Estimating covariance matrices in high-dimensional portfolios with nested and one-factor structures.
method Combining random matrix theory, free probability, deterministic equivalents, and two-step covariance estimators.
result Two-step estimators improve financial metrics in complex and one-factor covariance models.
The package High-dimensional Metrics (\Rpackage{hdm}) is an evolving collection of statistical methods for estimation and quantification of uncertainty in high-dimensional approximately sparse models. It focuses on providing confidence intervals and significance testing for (possibly many) low-dimensional subcomponents…
Simple model explains manifold structure in high-dimensional data.
problem Understanding manifold structure in high-dimensional data.
method Latent Metric Model with latent variables, correlation, and stationarity.
result Establishes statistical explanation for manifold hypothesis.
Compressing data helps learn Mahalanobis metrics effectively.
problem Learning Mahalanobis metrics in high-dimensional spaces.
method Randomly compress data to train a full-rank metric in a reduced feature space.
result Theoretical guarantees on error for Mahalanobis metric learning, independent of ambient dimension.
DADApy analyzes high-dimensional data manifolds in Python.
problem Analyzing complex, high-dimensional data.
method Estimating intrinsic dimension, density, clustering, comparing distance metrics.
result Effective analysis of data manifolds in Python.
Study reveals how high-dimensional models are vulnerable to consistent adversarial attacks.
problem Understanding the vulnerability of high-dimensional linear classifiers to adversarial attacks.
method Introducing a new error metric to quantify model vulnerability, and rigorously characterizing these metrics in asymptotic settings.
result As models become more overparameterized, their vulnerability to label-preserving perturbations increases.
Study on spaces of metrics with intermediate curvature bounds.
problem Understanding spaces of metrics with lower bounds on intermediate curvatures.
method Analyzing spaces of Riemannian metrics with specific curvature bounds on high-dimensional Spin-manifolds.
result Spaces of metrics with positive p-curvature and k-positive Ricci curvature have non-trivial homotopy groups.
A new perceptual adjustment query for metric learning reduces complexity in high-dimensional data.
problem Metric learning in high-dimensional data with limited human feedback.
method Inverted measurement scheme and two-stage estimator for PAQs.
result Sample complexity guarantees for the two-stage estimator of metric learning from PAQs.
The paper proves rigidity for mapping class group actions on metrics of positive scalar curvature.
problem Rigidity of mapping class group actions on metrics of positive scalar curvature.
method Parametrised Morse theory, 2-index theorem, sphere computations.
result Rigidity theorem for mapping class group action on positive scalar curvature metrics.
The paper develops methods for welfare analysis in dynamic models.
problem Estimating welfare metrics in complex, high-dimensional models.
method Dual and doubly robust representations, Lasso and Neural Network estimators.
result Automatic debiasing of welfare metrics without needing bias correction.
New method for high-dimensional linear regression using empirical Bayes.
problem Estimating prior in high-dimensional linear regression.
method Variational empirical Bayes approach with NPMLE and mean field approximation.
result Established asymptotic consistency and computational efficiency of the method.
Generative adversarial networks learn optimal transport maps efficiently.
problem Learning optimal transport maps between high-dimensional distributions.
method Proposed a GAN with discriminator objective as 2-Wasserstein metric.
result Generator learns optimal transport map during training.
In general, the clustering problem is NP-hard, and global optimality cannot be established for non-trivial instances. For high-dimensional data, distance-based methods for clustering or classification face an additional difficulty, the unreliability of distances in very high-dimensional spaces. We propose a distance-ba…
We prove that the space of complete, finite volume, pinched negatively curved Riemannian metrics on a smooth high-dimensional manifold is either empty or it is highly non-connected, provided their behavior at infinity is similar.
We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize some of the popular metrics, such as the Jaccard and bag distances on sets, Man…
Regularization improves logistic regression performance in high-dimensional settings.
problem Improving logistic regression in scenarios with many parameters and observations.
method Introducing a convex regularizer to the negative log-likelihood function to encourage desired structures.
result Explicit expressions for various performance metrics of regularized logistic regression are derived.
We study the Teichmüller space of negatively curved metrics on a high dimensional manifold, with applications to bundles with negatively curved fibers.
A novel GPUM constructs Gaussian Processes for unknown manifolds with probabilistic metrics.
problem High-dimensional data on unknown manifolds with non-Euclidean geometry.
method Bayesian Gaussian Processes latent variable models (BGPLVM), Riemannian geometry, probabilistic metric tensor, Brownian Motion.
result GPUM provides more accurate predictions on unknown manifolds compared to traditional methods.
New method makes quality metrics scale-invariant for high-dimensional data.
problem Scale sensitivity in quality metrics affects the accuracy of data projections.
method Analytical and empirical investigation of stress and KL divergence; introduction of a scale-invariant technique.
result The proposed technique accurately captures expected behavior and makes metrics scale-invariant.
New Einstein metrics found on manifolds with opposite curvature signs.
problem Finding Einstein metrics with opposite curvature signs on manifolds.
method Reviewing and extending previous work on high-dimensional smooth closed manifolds.
result Proved various related results, including new Einstein metrics.
In this paper, we address the problem of hidden common variables discovery from multimodal data sets of nonlinear high-dimensional observations. We present a metric based on local applications of canonical correlation analysis (CCA) and incorporate it in a kernel-based manifold learning technique.We show that this metr…
Proposes a new method for efficient manifold denoising robust to high dimensional noise.
problem Efficiently denoise manifolds in high dimensional spaces with complicated noise.
method Landmark diffusion and optimal shrinkage under high dimensional noise and compact manifold setup.
result Systematic comparison with other algorithms on simulated and real datasets shows superior performance.
New metrics avoid high-dimensional analysis challenges, proving convergence without 'curse of dimensionality'.
problem High-dimensional analysis challenges in empirical measure convergence.
method Proposed a new class of probability metrics free of the curse of dimensionality.
result Convergence of empirical measures is free of the curse of dimensionality.
IBPF algorithm tackles high-dimensional parameter learning for complex systems.
problem Learning high-dimensional parameters in complex, partially observed, and nonlinear systems.
method Iterated Block Particle Filter (IBPF) for graphical state space models.
result IBPF algorithm consistently beats the curse of dimensionality across various experiments.
A new tree-sliced Wasserstein distance improves optimal transport computations.
problem Computational and statistical drawbacks in optimal transport.
method Introducing tree metrics and averaging Wasserstein distances using random tree metrics.
result Tree-sliced Wasserstein distance outperforms other methods on benchmarks.
Quantum probability metrics improve distribution comparison in high dimensions.
problem Challenges in comparing probability distributions, especially in high-dimensional and non-compact domains.
method Quantum probability metrics (QPMs) derived from quantum state spaces, overcoming limitations of MMD.
result QPMs offer enhanced sensitivity to subtle distributional differences in high dimensions and improve performance in generative modeling.
Geometric framework detects outliers in high-dimensional data.
problem Detecting outliers in high-dimensional data.
method Geometric framework exploiting manifold structure.
result Significant improvement in outlier detection in high-dimensional data.
CAMEL enhances manifold embedding and learning with curvature metrics.
problem High-dimensional data classification, dimension reduction, and visualization.
method CAMEL uses a Riemannian manifold with curvature metrics for enhanced expressibility and interpretability.
result CAMEL outperforms state-of-the-art methods on high-dimensional datasets.
Proposes a neural network method to improve consistencies in high dimensional data analysis.
problem Inconsistencies among dimensionality reduction, clustering, and visualization tasks in high dimensional data analysis.
method Consistent Representation Learning (CRL) neural network that performs NLDR transformations to satisfy LGP constraints.
result Improves consistencies in data interpretation through end-to-end task execution.
Fine-grained visual categorization (FGVC) is to categorize objects into subordinate classes instead of basic classes. One major challenge in FGVC is the co-occurrence of two issues: 1) many subordinate classes are highly correlated and are difficult to distinguish, and 2) there exists the large intra-class variation (e…
A new IPM uses ReLU networks to measure probability discrepancies.
problem Measuring the difference between two probability distributions in high dimensions.
method Proposes a new parametric IPM using ReLU neural networks to optimize and distinguish between distributions.
result The proposed IPM has good convergence rates and can be used as a surrogate for other IPMs.
We use classical results in smoothing theory to extract information about the rational homotopy groups of the space of negatively curved metrics on a high dimensional manifold. It is also shown that smooth M-bundles over spheres equipped with fiberwise negatively curved metrics, represent elements of finite order in th…
We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal mapping, or scaling, of power law distributed data. Power law distributed data are foun…
Generic smooth boundaries for isoperimetric regions in 8D manifolds.
problem Understanding boundaries of isoperimetric regions in high-dimensional spaces.
method Generic regularity results for isoperimetric regions in closed Riemannian manifolds of dimension eight.
result Smooth nondegenerate boundaries for isoperimetric regions for generic metrics and volumes.
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.
New geometric analysis of PWSPDs balances density and geometry in high-dimensional data.
problem Balancing density and geometry in high-dimensional data.
method Power-weighted shortest-path distances (PWSPDs) and their geometric and computational analyses.
result High probability guarantees on the equivalence of PWSPDs on complete and nearest neighbor graphs.