CPCR mitigates bias in PCR for overparameterized models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new PCR method using SVD with sparse regularization.
Principal component regression (PCR) is a two-stage procedure that selects some principal components and then constructs a regression model regarding them as new explanatory variables. Note that the principal components are obtained from only explanatory variables and not considered with the response variable. To addre…
Principal Components Regression (PCR) is a traditional tool for dimension reduction in linear regression that has been both criticized and defended. One concern about PCR is that obtaining the leading principal components tends to be computationally demanding for large data sets. While random projections do not possess…
We solve principal component regression (PCR), up to a multiplicative accuracy , by reducing the problem to black-box calls of ridge regression. Therefore, our algorithm does not require any explicit construction of the top principal components, and is suitable for large-scale PCR instances. In…
VC-PCR improves prediction by clustering correlated variables.
This study analyzes prediction risk for PCR method in latent factor regression models.
Adaptive PCR improves panel data analysis with uniform guarantees.
PCR-LE achieves optimal rates for nonparametric regression over Sobolev spaces.
Principal component regression (PCR) is a widely used two-stage procedure: principal component analysis (PCA), followed by regression in which the selected principal components are regarded as new explanatory variables in the model. Note that PCA is based only on the explanatory variables, so the principal components a…
Improved fMRI analysis models enhance classification performance and select relevant brain regions.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
Principal component regression (PCR) is a simple, but powerful and ubiquitously utilized method. Its effectiveness is well established when the covariates exhibit low-rank structure. However, its ability to handle settings with noisy, missing, and mixed-valued, i.e., discrete and continuous, covariates is not understoo…
A new debiasing method for high-dimensional regression with applications to PCR.
A new method improves target selection for manipulating complex systems like the brain.
Consider a regression problem where there is no labeled data and the only observations are the predictions of experts over many samples . With no knowledge on the accuracy of the experts, is it still possible to accurately estimate the unknown responses ? Can one still detect the leas…
We analyze optimal weighted ridge regression in overparameterized linear models.
PLS-Lasso integrates dimension reduction into regression for financial index tracking.
New method speeds up learning of complex dynamical systems.
Software helps teach latent variable methods in multivariate data analytics.
PCHAL and PCHAR use principal components to speed up HAL and HAR methods.
We show how to efficiently project a vector onto the top principal components of a matrix, without explicitly computing these components. Specifically, we introduce an iterative algorithm that provably computes the projection using few calls to any black-box routine for ridge regression. By avoiding explicit principal …
Paper presents a faster classical algorithm for principal component regression.
Survey of SDR methods for high-dimensional regression and embedding.
New approach to quantify posterior concentration rates using Wasserstein dynamics.
We propose a new two stage algorithm LING for large scale regression problems. LING has the same risk as the well known Ridge Regression under the fixed design setting and can be computed much faster. Our experiments have shown that LING performs well in terms of both prediction accuracy and computational efficiency co…
Study on reducing dimensionality in high-dimensional regression with kernel methods and stability analysis.
Linear models can overfit without harming OOD generalization under certain conditions.
Efficient private matrix analysis algorithms for recent variants.
In the multiple linear regression setting, we propose a general framework, termed weighted orthogonal components regression (WOCR), which encompasses many known methods as special cases, including ridge regression and principal components regression. WOCR makes use of the monotonicity inherent in orthogonal components …
We propose a new method for supervised learning, especially suited to wide data where the number of features is much greater than the number of observations. The method combines the lasso () sparsity penalty with a quadratic penalty that shrinks the coefficient vector toward the leading principal components of …
PCA (Principal Component Analysis) and its variants areubiquitous techniques for matrix dimension reduction and reduced-dimensionlatent-factor extraction. One significant challenge in using PCA, is thechoice of the number of principal components. The information-theoreticMDL (Minimum Description Length) principle gives…
Two Kähler structures are PCR equivalent in the Siegel domain.
Proposes an online method for high-dimensional streaming data.
Study reduces financial dynamics complexity using PCA for NASDAQ, oil, gold, and USD.
We propose a penalized orthogonal-components regression (POCRE) for large p small n data. Orthogonal components are sequentially constructed to maximize, upon standardization, their correlation to the response residuals. A new penalization framework, implemented via empirical Bayes thresholding, is presented to effecti…
We study least squares linear regression over uncorrelated Gaussian features that are selected in order of decreasing variance. When the number of selected features is at most the sample size , the estimator under consideration coincides with the principal component regression estimator; when , the esti…
Technological change and innovation are vitally important, especially for high-tech companies. However, factors influencing their future research and development (R&D) trends are both complicated and various, leading it a quite difficult task to make technology tracing for high-tech companies. To this end, in this pape…
Noisy Pooled PCR tests large groups more efficiently.
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
Kernelized PCovR reveals structure-property relations in chemistry and materials.
Deep learning accelerators efficiently train over vast and growing amounts of data, placing a newfound burden on commodity networks and storage devices. A common approach to conserve bandwidth involves resizing or compressing data prior to training. We introduce Progressive Compressed Records (PCRs), a data format that…
We compare the risk of ridge regression to a simple variant of ordinary least squares, in which one simply projects the data onto a finite dimensional subspace (as specified by a Principal Component Analysis) and then performs an ordinary (un-regularized) least squares regression in this subspace. This note shows that …
Motivated by the Bagging Partial Least Squares (PLS) and Principal Component Analysis (PCA) algorithms, we propose a Principal Model Analysis (PMA) method in this paper. In the proposed PMA algorithm, the PCA and the PLS are combined. In the method, multiple PLS models are trained on sub-training sets, derived from the…
This paper introduces a new unsupervised method for dimensionality reduction via regression (DRR). The algorithm belongs to the family of invertible transforms that generalize Principal Component Analysis (PCA) by using curvilinear instead of linear features. DRR identifies the nonlinear features through multivariate r…
We explore the effect of past market movements on the instantaneous correlations between assets within the futures market. Quantifying this effect is of interest to estimate and manage the risk associated to portfolios of futures in a non-stationary context. We apply and extend a previously reported method called the P…
New methods combine low and high-fidelity data for accurate surrogate modeling.
Proposes a new method for multivariate functional regression.