Proposes an online method for high-dimensional streaming data.
problem Increasing variable dimensions with sample size in online kernel sliced inverse regression.
method Introduces approximate linear dependence condition and dictionary variable sets to address the problem. Transforms into online generalized eigen-decomposition problem and uses stochastic optimization for updates.
result Achieves close performance to batch processing kernel sliced inverse regression.
This paper proposes a novel kernel approach to linear dimension reduction for supervised learning. The purpose of the dimension reduction is to find directions in the input space to explain the output as effectively as possible. The proposed method uses an estimator for the gradient of regression function, based on the…
New method for reducing dimensions of distributional data.
problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.
Survey of SDR methods for high-dimensional regression and embedding.
problem Reducing dimensionality in high-dimensional data.
method Involves both statistical and machine learning approaches, covering inverse and forward regression methods.
result Supervised Kernel Dimension Reduction is equivalent to supervised PCA.
Optimizes differentially private kernel learning with random projection.
problem Privacy-preserving learning algorithms with optimal performance.
method Differentially private kernel ERM algorithm based on random projection in reproducing kernel Hilbert space.
result Achieves minimax-optimal excess risk rates for various loss functions.
A new deep neural network tackles nonlinear functional regression with improved dimensionality reduction.
problem Nonlinear functional regression in infinite-dimensional functional data analysis.
method Functional deep neural network with adaptive kernel embedding and projection steps.
result Explicit rates of approximating nonlinear smooth functionals are derived, and the network is shown to be effective in both simulated and real datasets.
Solves kernel dimension reduction while making features interpretable.
problem Making kernel dimension reduction methods interpretable.
method Projects onto a subspace before kernel feature mapping, using ISM for optimization.
result Extends ISM's theoretical guarantees to a family of kernels, enabling broader applicability.
In statistical learning, high covariate dimensionality poses challenges for robust prediction and inference. To address this challenge, supervised dimension reduction is often performed, where dependence on the outcome is maximized for a selected covariate subspace with smaller dimensionality. Prevalent dimension reduc…
Kernel PCA helps analyze multivariate extremes and clusters them effectively.
problem Analyzing the dependence structure of multivariate extremes.
method Kernel PCA as a method for clustering and dimension reduction.
result Kernel PCA preimages effectively identify clusters in multivariate extremes.
This paper improves HSIC-based dimensionality reduction for non-linear kernels.
problem Non-convexity of HSIC objective function for non-linear kernels limits optimization efficiency.
method Spectral optimization algorithm with local guarantees and principled initialization.
result Empirical improvements by a factor of 105 in runtime with lower errors. This study compares DR methods with kernel variations for face image analysis.
problem High dimensionality, noise, and correlation in data.
method Reviews and comparative study of PCA, LDA, KPCA, KLDA, SKPCA.
result SKPCA outperforms other methods in gender classification on face databases.
The purpose of sufficient dimension reduction (SDR) is to find the low-dimensional subspace of input features that is sufficient for predicting output values. In this paper, we propose a novel distribution-free SDR method called sufficient component analysis (SCA), which is computationally more efficient than existing …
Paper compares dimension reduction methods using topological analysis on EEG data.
problem Comparing dimension reduction methods on EEG data.
method Topological data analysis, including persistent homology, Wasserstein distance, and hypothesis tests.
result Different dimension reduction methods show significant qualitative differences across topological homologies.
Study shows how to effectively predict functions on manifolds using kernel methods.
problem Regression on manifolds with limited data.
method Reproducing kernel Hilbert space methods, Weyl law, effective dimension.
result Kernel regression estimator yields minimax-optimal error bounds controlled by effective dimension.
GDMaps reduces high-dimensional data to lower dimensions for better classification.
problem High-dimensional data classification and representation.
method Grassmannian Diffusion Maps technique for nonlinear dimensionality reduction.
result GDMaps effectively identifies intrinsic subspace structures in high-dimensional data.
New method circumvents curse of dimensionality in Laplacian estimation.
problem High-dimensional data challenges spectral clustering and diffusion maps.
method Kernelized Laplacian estimation via reproducing kernel Hilbert space.
result Non-asymptotic statistical rates show improved performance in high dimensions.
New bounds on KPCA efficiency reveal conditions for fast convergence.
problem Lack of theoretical understanding of KPCA efficiency.
method Lower and upper bounds on KPCA efficiency involving empirical eigenvalues and new variance quantities.
result Fast convergence rates achievable for certain kernels, highlighting dataset properties.
Study on reducing dimensionality in high-dimensional regression with kernel methods and stability analysis.
problem Analyzing errors in high-dimensional regression with dimensionality reduction and kernel regression.
method Derive a stability result for kernel regression with Wasserstein distance and apply it to PCA to deduce convergence rates.
result Two-step procedure yields useful convergence rates in semi-supervised settings.
A new geometry-preserving method for interpreting compositional data.
problem Statistical challenges in high-dimensional compositional data.
method Geometry-preserving framework for dimension reduction of compositional data.
result Identification of a central compositional subspace for compositional predictors.
We propose a method for feature selection that employs kernel-based measures of independence to find a subset of covariates that is maximally predictive of the response. Building on past work in kernel dimension reduction, we show how to perform feature selection via a constrained optimization problem involving the tra…
Develops a nonparametric graphical model for conditional independence.
problem Evaluation of conditional independence without distributional assumptions.
method Nonlinear sufficient dimension reduction techniques applied to a nonparametric graphical model.
result Method outperforms existing methods in non-Gaussian settings and high-dimensional data.
In this paper, we propose a novel supervised learning method that is called Deep Embedding Kernel (DEK). DEK combines the advantages of deep learning and kernel methods in a unified framework. More specifically, DEK is a learnable kernel represented by a newly designed deep architecture. Compared with pre-defined kerne…
New algorithms for clustering and dimension reduction using relative von Neumann entropy.
problem Clustering and dimension reduction for complex data sets.
method Construct graphs from data points, select graph maximizing relative von Neumann entropy, use eigenvectors for dimension reduction.
result Outperforms existing methods on non-trivial data sets.
Let (X,T1,0X) be a compact connected orientable CR manifold of dimension 2n+1 with non-degenerate Levi curvature. Assume that X admits a connected compact Lie group action G. Under certain natural assumptions about the group action G, we show that the G-invariant Szegö kernel for (0,q) forms is a comp…
New insights into tSNE for large datasets.
problem Limitations of tSNE in handling large datasets.
method Identified continuum limit of tSNE objective function, proposed rescaled model.
result Rescaled model has a consistent limit for large datasets.
Bayesian optimization (BO) has been broadly applied to computational expensive problems, but it is still challenging to extend BO to high dimensions. Existing works are usually under strict assumption of an additive or a linear embedding structure for objective functions. This paper directly introduces a supervised dim…
EnEMF uses Epanechnikov kernel for high-dimensional filtering, improving accuracy and robustness.
problem Suboptimal Gaussian mixture kernel density estimates in high-dimensional settings.
method Ensemble Epanechnikov mixture filter (EnEMF) using optimal Epanechnikov kernel.
result EnEMF reduces error per particle on high-dimensional systems like Lorenz '96.
A new algorithm interprets DR dimensions and selects features.
problem Lack of interpretability in DR algorithms.
method I-KDR algorithm that maps data to a lower dimensional space with interpretable dimensions and feature selection.
result I-KDR provides better interpretations and higher discriminative performance.
Paper develops KMS Wasserstein for high-dimensional data reduction.
problem Optimal transport's curse of dimensionality in high-dimensional data.
method Kernel max-sliced (KMS) Wasserstein distance for dimensionality reduction.
result Sharp finite-sample guarantees for KMS p-Wasserstein distance. A new method for fair representation learning using PLS.
problem Fairness in representation learning for data reduction.
method Proposes Fair Partial Least Squares (PLS) components with fairness constraints.
result The new method outperforms standard fair PCA methods on various datasets.
Overview of geometric analysis for manifold learning.
problem Analyzing high-dimensional data via spectral embeddings.
method Heat kernel and eigenfunctions on Riemannian manifolds.
result Uniform control of spectral embeddings on key classes of manifolds.
A method for reducing dimensions in Fréchet regression models.
problem Complex data objects in metric space-valued responses.
method Mapping metric-space valued random objects to real-valued variables and applying classical SDR.
result Consistent and asymptotically convergent method for Fréchet SDR.
Unified framework for spectral methods, kernel learning, and manifold unfolding.
problem Tackles the unification and optimization of spectral dimensionality reduction methods.
method Unified spectral methods as kernel PCA, kernel learning by SDP, and detailed explanation of MVU variants.
result Unified understanding and optimization of manifold learning techniques.
Survey of kernels, RKHS, and their applications in machine learning.
problem Understanding kernels and their applications in machine learning.
method Review of historical context, mathematical definitions, and practical applications of kernels.
result Comprehensive overview of kernels, RKHS, and their applications.
DM uses semigroup property to tune diffusion time for better data analysis.
problem Difficulty in tuning diffusion time for optimal data analysis.
method Proposes a semigroup criterion to select diffusion time.
result Effective and robust method for picking diffusion time.
We propose a representation of Gaussian processes (GPs) based on powers of the integral operator defined by a kernel function, we call these stochastic processes integral Gaussian processes (IGPs). Sample paths from IGPs are functions contained within the reproducing kernel Hilbert space (RKHS) defined by the kernel fu…
Identifies a gradient flow to solve kernel learning problems with noise reduction.
problem Kernel learning problem with Gaussian noise.
method Riemannian gradient flow with continuous Lyapunov functionals.
result Flow reduces noise and finds stationary points.
Study pure exploration in high-dimensional feature spaces using adaptive embeddings.
problem Overcoming the curse of dimensionality in pure exploration bandits.
method Adaptive embedding of feature representations into lower-dimensional spaces, carefully dealing with model misspecification.
result Sample complexity guarantees that depend on the effective dimension of feature spaces in kernel or neural representations.
Gaussian Process Latent Variable Model (GPLVM) is a flexible framework to handle uncertain inputs in Gaussian Processes (GPs) and incorporate GPs as components of larger graphical models. Nonetheless, the standard GPLVM variational inference approach is tractable only for a narrow family of kernel functions. The most p…
New method for spatiotemporal data regression using Gaussian processes.
problem Regression in spatiotemporal random fields.
method Empirical Bayes approach, tight Gaussian measures, truncation scheme.
result Effective dimension reduction through time-varying angular spectra.
New method corrects missing data bias in dimension reduction.
problem Missing data complicates high-dimensional data analysis.
method Developed a bias-corrected Gram matrix for heterogeneous missingness.
result Proposed method improves dimension reduction techniques significantly.
String kernels are attractive data analysis tools for analyzing string data. Among them, alignment kernels are known for their high prediction accuracies in string classifications when tested in combination with SVM in various applications. However, alignment kernels have a crucial drawback in that they scale poorly du…
New kernels allow learning from non-separable data.
problem Learning from non-separable data.
method Introducing entangled kernels and a two-step algorithm.
result Efficient algorithm for learning entangled kernels.
Sparse model for noisy datasets using hierarchical regularization.
problem Learning from large noisy datasets with sparse representations.
method Hierarchical learning strategy with projection-based penalty operators.
result Efficient sparse model reconstruction and generalizability on real datasets.
One of the major problems in natural language processing (NLP) is the word sense disambiguation (WSD) problem. It is the task of computationally identifying the right sense of a polysemous word based on its context. Resolving the WSD problem boosts the accuracy of many NLP focused algorithms such as text classification…
Dimensionality reduction is an important step in processing the hyperspectral images (HSI) to overcome the curse of dimensionality problem. Linear dimensionality reduction methods such as Independent component analysis (ICA) and Linear discriminant analysis (LDA) are commonly employed to reduce the dimensionality of HS…
Unified quadrature framework for large-scale kernel machines.
problem Efficiently approximating kernel functions for large-scale machine learning.
method Deterministic and randomized interpolatory rules for numerical integration of kernel functions.
result The proposed method reduces the number of nodes needed for accurate kernel approximation.
Thanks to their versatility, ease of deployment and high-performance, surrogate models have become staple tools in the arsenal of uncertainty quantification (UQ). From local interpolants to global spectral decompositions, surrogates are characterised by their ability to efficiently emulate complex computational models …