We prove the statistical consistency of kernel Partial Least Squares Regression applied to a bounded regression learning problem on a reproducing kernel Hilbert space. Partial Least Squares stands out of well-known classical approaches as e.g. Ridge Regression or Principal Components Regression, as it is not defined as…
Proposes a method for coarse graph alignment using sparse partial least squares.
problem Aligning graphs with community structures when there's no natural one-to-one mapping.
method Sparse partial least squares method incorporating observed graph structures and imposing sparsity.
result Demonstrates effectiveness in simulations.
This paper presents regression models obtained from a process of blind prediction of peptide binding affinity from provided descriptors for several distinct datasets as part of the 2006 Comparative Evaluation of Prediction Algorithms (COEPRA) contest. This paper finds that kernel partial least squares, a nonlinear part…
New algorithm extracts shared latent space for cortico-muscular interactions.
problem Challenges of high dimensionality and limited sample sizes in multivariate cortico-muscular analysis.
method Structured and sparse partial least squares coherence (ssPLSC) algorithm.
result ssPLSC achieves competitive or better performance in scenarios with limited sample sizes and high noise levels.
The derivation of statistical properties for Partial Least Squares regression can be a challenging task. The reason is that the construction of latent components from the predictor variables also depends on the response variable. While this typically leads to good performance and interpretable models in practice, it ma…
Functional PLS improves prediction and inference for scalar responses from functional predictors.
problem Estimating scalar responses from functional predictors in an ill-posed inverse problem.
method Functional partial least squares (PLS) estimator with adaptive early stopping and new tests.
result PLS attains nearly minimax-optimal convergence rates and detects local alternatives.
Improves Bayesian optimisation for engineering design problems with many variables.
problem Efficiently searching for global minima in high-dimensional design spaces.
method Integrates input and output data to identify a reduced latent subspace using probabilistic partial least squares.
result Significant improvements in convergence to the global minimum compared to existing methods.
A new PMA method combines PCA and PLS for better classification.
problem Improving classification performance in data analysis.
method Combines PCA and PLS through multiple sub-PLS models and PCA on the joint coefficient matrix.
result The proposed PMA method achieves better classification performance and stability.
Bayesian optimization selects wavelengths for sugar content estimation in NIR spectroscopy.
problem Improving prediction accuracy and interpretability of spectral data for sugar content estimation.
method Formulated as a binary black-box optimization problem, proposed method uses Bayesian optimization with a sparse quadratic surrogate model and Thompson sampling.
result Improves prediction accuracy of partial least squares regression and yields more consistent wavelength regions.
Unified multi-view learning framework using OPLS with regularization and deep extensions.
problem Improving multi-view learning for classification and feature extraction.
method Orthonormalized Partial Least Squares (OPLS) with regularization and deep extensions.
result Unified multi-view learning framework with improved performance.
Dual-sPLS improves feature selection and prediction in high-dimensional data.
problem Relating variables to a response in high-dimensional chemometric problems.
method Generalizes PLS1 algorithm with dual norm penalizations and a shrinking ratio parameter.
result Favorably compares to similar regression methods on simulated and real chemical data.
Combines BTEM and T-PLS for accurate spectral recovery and calibration.
problem Calibrating pure spectra of minority components in mixtures without prior knowledge.
method Band target entropy minimization (BTEM) and target partial least squares (T-PLS).
result Estimated amounts from BTEM-T-PLS similar to MCR-ALS on simple mixtures, superior on complex ones.
Study reveals limits of PLS in multi-modal learning with correlated signals.
problem Understanding PLS performance in multi-modal learning with correlated signals.
method Random matrix theory analysis of spiked cross-covariance models.
result Identifies SNR and correlation regimes where PLS fails to recover any signal.
Paper proposes a recursive PLS model for optimal response to security threats.
problem Optimal response to security threats after violations have occurred.
method Recursive Partial Least Squares (PLS) model with factorial analysis of security events.
result The model optimally estimates security administrators' responses to threats.
The paper analyzes PLS-SVD in high-dimensional data integration, revealing its strengths and limitations.
problem Understanding the behavior of PLS-SVD in high-dimensional data integration.
method Analysis using random matrix theory and singular value decomposition.
result PLS-SVD exhibits counter-intuitive or limiting behavior in certain regimes and outperforms PCA when detecting common latent subspace.
The runtime for Kernel Partial Least Squares (KPLS) to compute the fit is quadratic in the number of examples. However, the necessity of obtaining sensitivity measures as degrees of freedom for model selection or confidence intervals for more detailed analysis requires cubic runtime, and thus constitutes a computationa…
P3LS preserves privacy while integrating data across companies.
problem Privacy concerns in cross-organizational data exchange and integration.
method Privacy-preserving federated learning technique using SVD-based PLS and random masks.
result Improves prediction performance on process-related indicators.
Simple soft sensor models improve prediction accuracy with small moving windows.
problem Improving prediction accuracy in soft sensing processes with limited historical data.
method Five simple soft sensor methodologies with small moving windows were compared.
result Small moving window sizes led to the lowest prediction errors for all methods.
High-dimensional data common in genomics, proteomics, and chemometrics often contains complicated correlation structures. Recently, partial least squares (PLS) and Sparse PLS methods have gained attention in these areas as dimension reduction techniques in the context of supervised data analysis. We introduce a framewo…
This paper reviews and compares supervised linear dimension-reduction techniques.
problem Lack of information in the response during unsupervised PCA reduces predictive performance.
method Review and comparison of supervised linear dimension-reduction techniques.
result PLS and LSPCA consistently outperform other techniques in simulations.
Proposes a new method for joint sample and feature selection in multi-view data.
problem Cannot detect latent subsets of samples and remove outliers.
method Weighted Sparse Partial Least Squares (ℓ∞/ℓ0-wsPLS) method for joint sample and feature selection. result Developed globally convergent algorithm and iterative algorithms for multi-view data fusion.
PLS-Lasso integrates dimension reduction into regression for financial index tracking.
problem Dimension reduction and regression are traditionally treated separately in multivariate data analysis.
method PLS-Lasso integrates dimension reduction directly into the regression process, presenting two formulations: PLS-Lasso-v1 and PLS-Lasso-v2.
result PLS-Lasso-v1 and PLS-Lasso-v2 outperform Lasso in financial index tracking.
DPLS improves asset pricing by capturing non-linear risk factor structures.
problem Estimating asset pricing models with non-linear risk factor structures.
method Deep Partial Least Squares (DPLS) for dynamic and flexible factor modeling.
result DPLS models outperform linear models in asset pricing, capturing non-linear risk factor interactions.
In this paper we propose a computationally efficient algorithm for on-line variable selection in multivariate regression problems involving high dimensional data streams. The algorithm recursively extracts all the latent factors of a partial least squares solution and selects the most important variables for each facto…
Bayesian optimization reduces hyperparameters for mixed variable design problems.
problem Optimizing designs with a large number of mixed continuous, integer, and categorical variables.
method Adaptive dimension reduction using partial least squares for fewer hyperparameters.
result Significant improvement in performance compared to genetic algorithms.
This paper reviews SDR methods for multivariate response regression.
problem Handling sufficient dimension reduction for multivariate response regression.
method Characterizes SDR estimators as inverse or forward regression methods.
result Pooled marginal, projective resampling, distance-based, ordinary least squares, partial least squares, and semiparametric SDR estimators are discussed.
Gradient-enhanced kriging reduces function evaluations for high-dimensional problems.
problem High-dimensional function evaluations are computationally expensive.
method Developed a new gradient-enhanced surrogate model using partial-least squares to reduce hyperparameters and correlation matrix size.
result Significantly reduces the number of function evaluations required for accurate surrogate models.
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
problem Analyzing the difference between PLS and OLS regression in terms of eigenvalue distributions.
method Examined the distance between PLS and OLS regression coefficients using the Mahalanobis distance and eigenvalue distributions of the regressor covariance matrix.
result Provided a bound on the distance between PLS and OLS regression coefficients that depends only on the eigenvalue distribution of the regressor covariance matrix.
A new method for CT using graph-based regularization.
problem Transfer calibrations between instruments without suitable transfer standards.
method Employing manifold regularization of PLS objective to enforce invariant projections in latent variable space.
result Implicit removal of inter-device variation in predictive directions.
Study shows time-varying stock returns across economic states.
problem Equity premium predictability varies by economic state.
method State-switching predictive regression using yield curve slope.
result The Aligned Economic Index improves stock return prediction.
Exponential family extensions of principal component analysis (EPCA) have received a considerable amount of attention in recent years, demonstrating the growing need for basic modeling tools that do not assume the squared loss or Gaussian distribution. We extend the EPCA model toolbox by presenting the first exponentia…
Paper improves gas species identification in complex mixtures using neural networks.
problem Identifying gas species in multi-gas mixtures with high accuracy.
method Multi-label neural networks with optimal thresholding for IR spectroscopy.
result Optimal thresholding improves classification performance over conventional methods.
edPLS adds Gaussian noise to PLS regression to protect data privacy.
problem Protecting sensitive data in PLS regression models.
method Integrates Gaussian noise into PLS algorithm based on global sensitivity.
result Effective at preserving privacy while maintaining competitive prediction accuracy.
Unified CCA methods for large-scale data with fast SGD algorithms.
problem Computational infeasibility of classical CCA methods for large-scale data.
method Unconstrained objective, stochastic gradient descent (SGD) algorithms.
result Significantly faster convergence and higher correlations than previous methods.
A new framework for PPLS combines noise estimation, optimization, and calibration.
problem Probabilistic PLS models need interpretable latent factors and calibrated uncertainty.
method End-to-end pipeline combining noise estimation, constrained optimization, and prediction calibration.
result Achieves near-nominal coverage and native calibrated uncertainty across benchmarks.
We prove rates of convergence in the statistical sense for kernel-based least squares regression using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is directly related to Kernel Partial Least Squares, a regression method that combines supervised dim…
R-PLS improves analysis of brain functional connectivity matrices.
problem Improving analysis of functional connectivity matrices in brain imaging.
method Introducing R-PLS, a generalization of PLS for symmetric positive definite matrices.
result R-PLS identifies key functional connections in brain imaging datasets.
We prove statistical rates of convergence for kernel-based least squares regression from i.i.d. data using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is related to Kernel Partial Least Squares, a regression method that combines supervised dimensio…
Word embeddings have been shown to be useful across state-of-the-art systems in many natural language processing tasks, ranging from question answering systems to dependency parsing. (Herbelot and Vecchi, 2015) explored word embeddings and their utility for modeling language semantics. In particular, they presented an …
Principal Component Analysis (PCA) is a very successful dimensionality reduction technique, widely used in predictive modeling. A key factor in its widespread use in this domain is the fact that the projection of a dataset onto its first K principal components minimizes the sum of squared errors between the original …
A new method for fair representation learning using PLS.
problem Fairness in representation learning for data reduction.
method Proposes Fair Partial Least Squares (PLS) components with fairness constraints.
result The new method outperforms standard fair PCA methods on various datasets.
Developed AI models for multi-gas detection in near IR spectrums.
problem Detecting multiple gases in near IR spectrums.
method Used Monte Carlo KNN and multi-resolution CNN, synthesized near IR spectrums, optimized kernel sizes and channels.
result Multi-resolution CNN outperforms other models.
Matrix factorizations and their extensions to tensor factorizations and decompositions have become prominent techniques for linear and multilinear blind source separation (BSS), especially multiway Independent Component Analysis (ICA), NonnegativeMatrix and Tensor Factorization (NMF/NTF), Smooth Component Analysis (Smo…
Partial Least Squares (PLS) methods have been heavily exploited to analyse the association between two blocs of data. These powerful approaches can be applied to data sets where the number of variables is greater than the number of observations and in presence of high collinearity between variables. Different sparse ve…
Deep learning is a form of machine learning for nonlinear high dimensional pattern matching and prediction. By taking a Bayesian probabilistic perspective, we provide a number of insights into more efficient algorithms for optimisation and hyper-parameter tuning. Traditional high-dimensional data reduction techniques, …
Many modern data mining applications are concerned with the analysis of datasets in which the observations are described by paired high-dimensional vectorial representations or "views". Some typical examples can be found in web mining and genomics applications. In this article we present an algorithm for data clusterin…
We investigate the optimal structure of dynamic regression models used in multivariate time series prediction and propose a scheme to form the lagged variable structure called Backward-in-Time Selection (BTS) that takes into account feedback and multi-collinearity, often present in multivariate time series. We compare …
We briefly review recent progress in techniques for modeling and analyzing hyperspectral images and movies, in particular for detecting plumes of both known and unknown chemicals. For detecting chemicals of known spectrum, we extend the technique of using a single subspace for modeling the background to a "mixture of s…