New framework tackles high-dimensional reliability analysis using surrogate models and active subspaces.
problem High computational cost and curse of dimensionality in reliability analysis of high-dimensional systems.
method Sparse Active Subspace (SAS) algorithm for identifying low-dimensional manifolds and constructing efficient surrogate models.
result Proposed framework significantly improves accuracy and efficiency of reliability analysis compared to existing methods.
Improved LDA method for better classification and dimensionality reduction.
problem Improving linear discriminant analysis for better classification performance.
method Integrates spectrally-corrected covariance matrix and regularized discriminant analysis.
result SRLDA has a linear classification global optimal solution under spiked model assumption.
The fundamental purpose of the present research article is to introduce the basic principles of Dimensional Analysis in the context of the neoclassical economic theory, in order to apply such principles to the fundamental relations that underlay most models of economic growth. In particular, basic instruments from Dime…
PCA simplifies multivariate extreme data analysis.
problem Analyzing multivariate extreme values with high-dimensional data.
method Principal Component Analysis (PCA) for dimensionality reduction.
result PCA helps preserve essential information for extreme value analysis.
The paper introduces a method for interpretable principal component analysis of high-dimensional time series.
problem Inconsistent and difficult-to-interpret principal component estimates in high-dimensional regimes.
method Localized sparse principal component analysis of spectral density matrices in frequency domain.
result Efficient algorithm for sparse-localized estimates of principal subspaces.
Interactive DR framework for comparing datasets.
problem Limited flexibility in existing DR methods for comparative analysis.
method Unified linear comparative analysis (ULCA) with interactive optimization and visualization.
result ULCA and optimization algorithm improve comparative analysis efficiency and flexibility.
Linear dimensionality reduction methods are a cornerstone of analyzing high dimensional data, due to their simple geometric interpretations and typically attractive computational properties. These methods capture many data features of interest, such as covariance, dynamical structure, correlation between data sets, inp…
Survey of factor analysis, PCA, variational inference, and VAE.
problem Dimensionality reduction and generative modeling of data.
method Variational inference, factor analysis, probabilistic PCA, and VAE.
result Derivation and explanation of ELBO, EM, and closed-form solutions.
We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal mapping, or scaling, of power law distributed data. Power law distributed data are foun…
The paper assesses dimensionality reduction for cryptocurrency link prediction.
problem Establishing a link between cryptocurrencies using dimensionality reduction techniques.
method Used canonical correlation analysis and principal component analysis on log returns and covariates of Bitcoin and Ethereum.
result Performance of dimensionality reduction techniques in forecasting Ethereum returns with Bitcoin features.
High dimensional data analysis is known to be as a challenging problem. In this article, we give a theoretical analysis of high dimensional classification of Gaussian data which relies on a geometrical analysis of the error measure. It links a problem of classification with a problem of nonparametric regression. We giv…
New method visualizes noisy data better than existing techniques.
problem Noisy data impairs data visualization methods.
method Functional Information Geometry (FIG) adapts EIG framework using functional data analysis.
result FIG outperforms EIG variant in capturing true structure, robustness, and speed.
We consider the high-dimensional discriminant analysis problem. For this problem, different methods have been proposed and justified by establishing exact convergence rates for the classification risk, as well as the l2 convergence results to the discriminative rule. However, sharp theoretical analysis for the variable…
Research covers geometry, analysis, and integration on infinite-dimensional spaces.
problem Exploring geometric and analytical structures in infinite-dimensional settings.
method Analyzes numerical schemes, Lie groups, connections, and integration theory.
result Developed new methods for integration and analysis on infinite-dimensional manifolds.
This work connects LLE, factor analysis, and probabilistic PCA through a stochastic perspective.
problem Exploring the theoretical connection between LLE, factor analysis, and probabilistic PCA.
method Solving the stochastic linear reconstruction of LLE using expectation maximization.
result LLE, factor analysis, and probabilistic PCA are shown to be connected through a stochastic perspective.
New methods integrate nonlinear, sparse, and multi-view aspects for high-dimensional data analysis.
problem Integrating nonlinear dependence, sparsity, and multi-view data in high-dimensional datasets.
method Proposes HSIC-SGCCA, SA-KGCCA, and TS-KGCCA methods for multi-view high-dimensional data analysis.
result HSIC-SGCCA outperforms competing methods in multi-view variable selection.
New method for tensor classification with missing data.
problem Handling incomplete tensor data in high-dimensional classification.
method High-dimensional tensor linear discriminant analysis with TGMM and Tensor LDA-MD.
result Established convergence rates and minimax optimal bounds for misclassification rate.
SEDA improves RLDA for high-dimensional data.
problem Inconsistent performance of RLDA in high-dimensional scenarios.
method Developed a non-asymptotic approximation of misclassification rate, derived new theoretical results on eigenvectors, and proposed SEDA algorithm.
result SEDA achieves higher classification accuracy and dimensionality reduction compared to existing LDA methods.
MediEncoder learns nonlinear representations for causal mediation analysis.
problem High-dimensional noisy covariates and mediators in biomedical studies.
method Coupled encoder-decoder architecture with cross-factor network.
result Improves estimation accuracy in high-dimensional causal mediation analysis.
New method analyzes complex multivariate pathways in high-dimensional data.
problem High-dimensional mediation analysis of multivariate exposures, mediators, and outcomes.
method Simultaneous variable selection, indirect effect matrix estimation, and prediction of multivariate outcomes.
result Identifies biologically interpretable genetic-neural-cognitive pathways.
PERCEPT detects changes in high-dimensional data streams using topological data analysis.
problem Detecting changes in high-dimensional data streams, especially when embedded in a low-dimensional space.
method Leverages topological data analysis to learn embedded topology as a point cloud via persistence diagrams, then applies non-parametric monitoring for detecting changes.
result Demonstrates efficient detection of online changes from high-dimensional data streams.
Study on reducing dimensionality in high-dimensional regression with kernel methods and stability analysis.
problem Analyzing errors in high-dimensional regression with dimensionality reduction and kernel regression.
method Derive a stability result for kernel regression with Wasserstein distance and apply it to PCA to deduce convergence rates.
result Two-step procedure yields useful convergence rates in semi-supervised settings.
Non-Gaussian component analysis (NGCA) is an unsupervised linear dimension reduction method that extracts low-dimensional non-Gaussian "signals" from high-dimensional data contaminated with Gaussian noise. NGCA can be regarded as a generalization of projection pursuit (PP) and independent component analysis (ICA) to mu…
Tensor Regression tackles high-dimensional data analysis.
problem Challenges in traditional data representation methods for high-dimensional data.
method Systematic study and analysis of tensor-based regression models.
result Provides solutions for specific regression tasks with multiway data.
Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.
problem Inherent computational complexity of archetypal analysis limits its practical applicability.
method Two preprocessing techniques: dimensionality reduction and representation cardinality reduction, using probabilistic geometry.
result The method effectively reduces scaling and provides near-optimal solutions for prediction errors.
The paper shows that ancient noncollapsed mean curvature flows have a blowdown of at most n-2 dimensions.
problem Understanding the blowdown of ancient noncollapsed mean curvature flows.
method Fine cylindrical analysis and fine neck analysis generalization.
result The blowdown of ancient noncollapsed mean curvature flows is at most n-2 dimensional.
Novel method converts time series data into functional data for high dimensional classification.
problem Small sample size problem in high dimensional time series data.
method Classwise Functional Principal Component Analysis (PCA) followed by Bayesian linear classifier.
result Demonstrated efficacy on synthetic and real data sets.
Paper formalizes multi-dimensional FSD using geometric methods.
problem Complex measure theory and calculus barriers to formalization in proof assistants.
method Geometric framework for first-order stochastic dominance in N dimensions.
result Geometric approach bypasses complex integration theory for direct comparison of survival probabilities.
Constructs Gabor frames for curved manifolds to detect boundaries.
problem Signal analysis on curved manifolds with boundaries.
method Higher-dimensional Gabor frames for local linearizations.
result Detection of higher-dimensional boundaries in curved manifolds.
Shape analysis and compuational anatomy both make use of sophisticated tools from infinite-dimensional differential manifolds and Riemannian geometry on spaces of functions. While comprehensive references for the mathematical foundations exist, it is sometimes difficult to gain an overview how differential geometry and…
New KNN test improves association analysis of high-dimensional sequencing data.
problem Challenges in using neural networks for high-dimensional sequencing data analysis.
method Kernel-based neural network (KNN) test for complex association analysis.
result KNN test outperforms SKAT in detecting non-linear and interaction effects.
This paper proposes a probabilistic neural network developed on the basis of time-series discriminant component analysis (TSDCA) that can be used to classify high-dimensional time-series patterns. TSDCA involves the compression of high-dimensional time series into a lower-dimensional space using a set of orthogonal tra…
The widely used genetic pleiotropic analysis of multiple phenotypes are often designed for examining the relationship between common variants and a few phenotypes. They are not suited for both high dimensional phenotypes and high dimensional genotype (next-generation sequencing) data. To overcome these limitations, we …
A hierarchical approach improves classification accuracy in large datasets.
problem Improving classification accuracy in large datasets with high dimensionality.
method Hierarchical subspace learning to scale manifold learning methods.
result Average 5% increase in classification accuracy.
New method explains high-dimensional text classifiers.
problem Limited explainability tools for high-dimensional inputs and neural networks.
method Theoretical high-dimensional properties in neural networks.
result Improved explainability for neural network classifiers.
A method for making machine learning units-equivariant using dimensional analysis.
problem Ensuring machine learning models respect dimensional consistency.
method Constructing dimensionless inputs and applying equivariant machine learning methods.
result Improved accuracy in tasks requiring dimensional consistency.
In this work, we develop a novel principal component analysis (PCA) for semimartingales by introducing a suitable spectral analysis for the quadratic variation operator. Motivated by high-dimensional complex systems typically found in interest rate markets, we investigate correlation in high-dimensional high-frequency …
Flow cytometry is often used to characterize the malignant cells in leukemia and lymphoma patients, traced to the level of the individual cell. Typically, flow cytometric data analysis is performed through a series of 2-dimensional projections onto the axes of the data set. Through the years, clinicians have determined…
Bayesian tree ensemble model for estimating treatment effects in high-dimensional survival data.
problem Estimating heterogeneous treatment effects in censored survival data with many covariates.
method Developed a Bayesian tree ensemble model with a horseshoe prior for adaptive shrinkage.
result Accurately estimates treatment effects in high-dimensional covariate spaces and non-linear functions.
Unified framework for high-dimensional bandit problems with low-dimensional structures.
problem Stochastic high-dimensional bandit problems with low-dimensional structures.
method Proposed a simple unified algorithm and a general analysis framework for the regret upper bound.
result Unified algorithm achieves comparable regret bounds in various high-dimensional bandit problems.
Overview of high-dimensional time series regression methods.
problem Estimation and inference with high-dimensional time series data.
method Limit theory for high-dimensional dependent data, asymptotic theory for time series regression, statistical learning methods.
result Main limit theory results and asymptotic theory for high-dimensional time series regression.
In recent years, the spectral analysis of appropriately defined kernel matrices has emerged as a principled way to extract the low-dimensional structure often prevalent in high-dimensional data. Here we provide an introduction to spectral methods for linear and nonlinear dimension reduction, emphasizing ways to overcom…
A novel method extracts topological features from word embeddings for text classification.
problem High dimensional and noisy text representations in natural language processing.
method Persistent homology for topological data analysis on word embeddings.
result Topological features outperform conventional text mining features on long textual documents.
dtSNE preserves local densities in low-dimensional embeddings.
problem Local density differences are not accurately preserved in tSNE and UMAP.
method dtSNE, which approximately conserves local densities.
result dtSNE provides more accurate local density depictions.
Mapper and Ball Mapper tools for complex data analysis.
problem Exploring and visualizing high-dimensional data and scalar functions.
method Combining Mapper and Ball Mapper, adding new features for encoding structure and symmetries.
result A new hybrid algorithm, Mapper on Ball Mapper, for comparing high-dimensional data descriptors.
Class prediction is an important application of microarray gene expression data analysis. The high-dimensionality of microarray data, where number of genes (variables) is very large compared to the number of samples (obser- vations), makes the application of many prediction techniques (e.g., logistic regression, discri…
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
New metrics on curve spaces improve shape analysis.
problem Discretization of curve spaces and metric completeness.
method Sobolev metrics on discrete regular curves, completeness analysis.
result The finite-dimensional Riemannian manifolds are complete.