A novel deep learning method for chemometric data improves performance over transfer learning.
problem Training deep neural networks from chemometric data with varying input sizes.
method Weight sharing in deep convolutional neural networks trained on multiple data sets of different sizes.
result Superior performance compared to transfer learning, especially when training on medium and small data sets.
We propose novel deep learning based chemometric data analysis technique. We trained L2 regularized sparse autoencoder end-to-end for reducing the size of the feature vector to handle the classic problem of the curse of dimensionality in chemometric data analysis. We introduce a novel technique of automatic selection o…
CNNs outperform standard chemometric methods for spectral data classification.
problem Reducing the need for pre-processing steps in spectral data analysis.
method Convolutional neural networks (CNNs) compared with SVMs and PLSR for classification and regression of spectral data.
result CNNs outperform standard chemometric methods, especially for classification tasks.
Tensor analysis tackles complex multidimensional data across fields.
problem Efficiently extracting information from high-dimensional data.
method Interdisciplinary approach combining statistics, optimization, and numerical linear algebra.
result Significant progress in tensor analysis over the last decade.
Dual-sPLS improves feature selection and prediction in high-dimensional data.
problem Relating variables to a response in high-dimensional chemometric problems.
method Generalizes PLS1 algorithm with dual norm penalizations and a shrinking ratio parameter.
result Favorably compares to similar regression methods on simulated and real chemical data.
This work tackles sparse coding in DLRA for interpretable multiway data.
problem Sparse coding in DLRA for interpretable multiway data.
method Proposes a new sparse-coding subproblem (MSC) and several algorithms to solve it.
result DLRA extends low-rank approximations, reducing variance and enhancing interpretability.
Modeling variability in tensor decomposition methods is one of the challenges of source separation. One possible solution to account for variations from one data set to another, jointly analysed, is to resort to the PARAFAC2 model. However, so far imposing constraints on the mode with variability has not been possible.…
A new method constrains PARAFAC2 for better pattern recovery.
problem Challenges in analyzing multi-way measurements with variations across one mode.
method AO-ADMM approach to fit PARAFAC2 model with flexible constraints.
result The proposed method allows for flexible constraints, recovers patterns accurately, and is computationally efficient.
The paper explores partial identifiability in nonnegative matrix factorization under specific conditions.
problem Identifying specific columns of the matrices in nonnegative matrix factorization.
method Mathematical rigor and geometric interpretation to analyze partial identifiability of columns in nonnegative matrix factorization.
result The partial uniqueness of a single column of C or S can be guaranteed under certain sparsity and algebraic conditions. RAMANMETRIX simplifies Raman spectroscopy data analysis.
problem Complexity of Raman spectroscopy data processing.
method User-friendly software with chemometric analysis workflow.
result Facilitates integration of Raman spectroscopy into clinical routines.
High-dimensional data common in genomics, proteomics, and chemometrics often contains complicated correlation structures. Recently, partial least squares (PLS) and Sparse PLS methods have gained attention in these areas as dimension reduction techniques in the context of supervised data analysis. We introduce a framewo…
The PARAFAC tensor decomposition has enjoyed an increasing success in exploratory multi-aspect data mining scenarios. A major challenge remains the estimation of the number of latent factors (i.e., the rank) of the decomposition, which yields high-quality, interpretable results. Previously, we have proposed an automate…
High-dimensional tensors or multi-way data are becoming prevalent in areas such as biomedical imaging, chemometrics, networking and bibliometrics. Traditional approaches to finding lower dimensional representations of tensor data include flattening the data and applying matrix factorizations such as principal component…
Novel method converts time series data into functional data for high dimensional classification.
problem Small sample size problem in high dimensional time series data.
method Classwise Functional Principal Component Analysis (PCA) followed by Bayesian linear classifier.
result Demonstrated efficacy on synthetic and real data sets.
New methods explain NE embeddings by identifying key variables.
problem Lack of interpretability in NE techniques.
method Combining PCA, Q-residuals, Hotelling's T2, and visualization.
result Identifies discriminatory features not seen in standard approaches.
New ADMM method for PARAFAC2 tensor decomposition with flexible regularization.
problem Challenges in applying regularisation to the evolving mode of PARAFAC2.
method Alternating Direction Method of Multipliers (AO-ADMM) for PARAFAC2 tensor fitting.
result The proposed ADMM-based approach accurately recovers underlying components from simulated data.
Tensor decomposition is an important technique for capturing the high-order interactions among multiway data. Multi-linear tensor composition methods, such as the Tucker decomposition and the CANDECOMP/PARAFAC (CP), assume that the complex interactions among objects are multi-linear, and are thus insufficient to repres…
Tucker decomposition is the cornerstone of modern machine learning on tensorial data analysis, which have attracted considerable attention for multiway feature extraction, compressive sensing, and tensor completion. The most challenging problem is related to determination of model complexity (i.e., multilinear rank), e…
A new method for CT using graph-based regularization.
problem Transfer calibrations between instruments without suitable transfer standards.
method Employing manifold regularization of PLS objective to enforce invariant projections in latent variable space.
result Implicit removal of inter-device variation in predictive directions.
The paper analyzes an ensemble of randomly projected linear discriminants for high-dimensional data.
problem Classification issues in small samples of high-dimensional data.
method Asymptotic analysis using random matrix theory.
result The ensemble offers a performance advantage under certain conditions.
Proposes a new model to handle latent structure methods.
problem Lack of model inference, generative form, and unidentifiable parameters in latent structure methods.
method Generative Flexible Latent Structure Regression (GFLSR) model structure.
result Proposed model allows for model inference and leads to potential probabilistic model.
We use partial class memberships in soft classification to model uncertain labelling and mixtures of classes. Partial class memberships are not restricted to predictions, but may also occur in reference labels (ground truth, gold standard diagnosis) for training and validation data. Classifier performance is usually ex…
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
problem Analyzing the difference between PLS and OLS regression in terms of eigenvalue distributions.
method Examined the distance between PLS and OLS regression coefficients using the Mahalanobis distance and eigenvalue distributions of the regressor covariance matrix.
result Provided a bound on the distance between PLS and OLS regression coefficients that depends only on the eigenvalue distribution of the regressor covariance matrix.
A new method selects regions of interest in GC-MS data without prior target selection.
problem Challenges in GC-MS data analysis due to fragmentation and shared fragment ions.
method Uses a pseudo F-ratio moving window (ψFRMV) to automatically select regions of interest. result Algorithm can accurately identify signal regions in GC-MS data.