We compare the risk of ridge regression to a simple variant of ordinary least squares, in which one simply projects the data onto a finite dimensional subspace (as specified by a Principal Component Analysis) and then performs an ordinary (un-regularized) least squares regression in this subspace. This note shows that …
Reduced-rank method improves least-squares regression under output regularity.
problem Least-squares regression with infinite dimensional outputs.
method Reduced-rank method for solving least-squares problems with output regularity assumptions.
result Learning bounds and improved statistical performance compared to full-rank method.
Correcting bias in least squares regression with volume-rescaled sampling.
problem Bias in linear least squares solutions without distributional assumptions.
method Volume-rescaled sampling to correct bias in i.i.d. samples.
result Combined sample becomes unbiased with rescaled volume additional sample.
Local control regression improves portfolio optimization accuracy.
problem Expensive and inaccurate global control regression for portfolio optimization.
method Introduced local control regression combined with adaptive grids.
result Choosing a coarse grid for local regression produces accurate results.
We find a convex model for traditional nonlinear regression under L2 loss.
problem Nonlinear regression under L2 loss with non-convex optimization.
method Showed a convex nonlinear regression model for least squares problem.
result Existence of a convex model simplifies training complex systems.
We prove the statistical consistency of kernel Partial Least Squares Regression applied to a bounded regression learning problem on a reproducing kernel Hilbert space. Partial Least Squares stands out of well-known classical approaches as e.g. Ridge Regression or Principal Components Regression, as it is not defined as…
The study explores nonparametric regression with shape constraints using least squares estimation.
problem Nonparametric regression under shape constraints.
method Least squares estimation (LSE) with focus on isotonic, unimodal, convex, and additive shape-restricted regression.
result Adaptive nature of the LSE and its risk behavior, with pointwise limiting distribution theory for isotonic regression.
Optimal rates for sketched-regularized algorithms in least-squares regression over Hilbert spaces.
problem Least-squares regression problem over Hilbert spaces with regularization.
method Combining regularized algorithms with projection, using randomized sketches and Nyström methods.
result Optimal rates for sketched-regularized algorithms with sketch dimension proportional to effective dimension.
Improved robustness in kernel-based regression via novel loss function and IRLS.
problem Noise sensitivity in kernel-based regression methods.
method Proposed ℓs-loss function and iteratively reweighted least squares (IRLS) optimization. result Improved noise robustness in kernel-based regression methods.
New inequality for regression risk with random design and noise.
problem Excess risk in least-squares regression with random design and heteroscedastic noise.
method Proved a new concentration inequality for the excess risk in least-squares regression with random design and heteroscedastic noise, separating linearized and quadratic processes.
result Generalized the approach to quadratic contrasts and random design.
Study improves least squares estimation for heavy-tailed errors.
problem Improving least squares estimation under heteroscedastic and heavy-tailed errors.
method Analyzes the rate of convergence of least squares estimator under bounded conditional variance and finitely many moments of errors.
result Upper bounds on rates of convergence of LSE for heavy-tailed errors are found.
New deep learning solver for high-dimensional derivative pricing.
problem High-dimensional derivatives pricing problems.
method Combines deep learning with least square regression for backward SDE solving.
result Accurate and efficient pricing of complex derivatives.
Improved robust regression for heavy-tailed and contaminated data.
problem Linear regression with heavy-tailed and adversarially contaminated covariates and responses.
method Applying a filtering algorithm to covariates and then using Huber regression, least trimmed squares, or least absolute deviation estimators on the remaining data.
result Near-optimal error rates achieved for the Huber regression estimator.
Study reveals KLMS algorithms as simplified GP regression models.
problem Understanding KLMS algorithms and their performance differences.
method Examined the relationship between online Gaussian process regression and KLMS algorithms.
result KLMS algorithms correspond to specific cases of a parametric model of posterior covariance.
Gradient flow in least squares regression is at least 1.69 times riskier than ridge regression.
problem Comparing the risk of gradient descent iterates to ridge regression in least squares regression.
method Continuous-time view of gradient descent, proving risk bounds.
result Gradient flow's risk is at least 1.69 times that of ridge regression.
New algorithm reduces error in regression problems.
problem Minimizing composite objective functions with quadratic and convex components.
method Stochastic dual averaging with constant step-size, proving convergence rate O(1/n).
result Extends least-squares regression to various convex regularizers and geometries.
We prove rates of convergence in the statistical sense for kernel-based least squares regression using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is directly related to Kernel Partial Least Squares, a regression method that combines supervised dim…
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
problem Analyzing the difference between PLS and OLS regression in terms of eigenvalue distributions.
method Examined the distance between PLS and OLS regression coefficients using the Mahalanobis distance and eigenvalue distributions of the regressor covariance matrix.
result Provided a bound on the distance between PLS and OLS regression coefficients that depends only on the eigenvalue distribution of the regressor covariance matrix.
R2T hybrid model improves robust regression for asymmetric noise.
problem Least-squares regression fails with asymmetric structured noise.
method Transformer encoder, compression NN, fixed symbolic equation.
result Median regression MSE of 6e-6 to 3.5e-5 on synthetic data.
Sparse linear regression, which entails finding a sparse solution to an underdetermined system of linear equations, can formally be expressed as an l0-constrained least-squares problem. The Orthogonal Least-Squares (OLS) algorithm sequentially selects the features (i.e., columns of the coefficient matrix) to greedil…
This paper reviews SDR methods for multivariate response regression.
problem Handling sufficient dimension reduction for multivariate response regression.
method Characterizes SDR estimators as inverse or forward regression methods.
result Pooled marginal, projective resampling, distance-based, ordinary least squares, partial least squares, and semiparametric SDR estimators are discussed.
Paper shows how to fit regression models on encrypted data.
problem Statistical analysis on encrypted data.
method Coordinate and accelerated gradient descent algorithms using FHE.
result Gradient descent outperforms in encrypted computational speed.
The paper examines prediction and estimation risks of ridgeless least squares under general error assumptions.
problem Prediction and estimation risks of ridgeless least squares under realistic error structures.
method Analysis of prediction and estimation risks under general regression error assumptions, including clustered or serial dependence.
result The benefits of overparameterization extend to time series, panel, and grouped data.
ESNs trained with Tikhonov least squares approximate ergodic dynamical systems in L2(μ) norm.
problem Approximating ergodic dynamical systems using ESNs.
method Tikhonov least squares regression on ESNs trained on observations from an ergodic dynamical system.
result ESNs trained with Tikhonov least squares approximate the target function in the L2(μ) norm.
This paper compares LSM and ANN/GBM for pricing American put options under a complex model.
problem Pricing American put options using advanced techniques.
method Least-Squares Monte Carlo (LSM) and Artificial Neural Network (ANN) and Gradient Boosted Machine (GBM) Trees.
result LSM outperforms ANN and GBM in pricing American put options.
Lecture notes on advanced linear regression methods.
problem Understanding the properties of linear regression estimators in high dimensions.
method Proposition-proof exploration of least squares, ridgeless, ridge, and lasso estimators.
result Detailed analysis of the existence, uniqueness, relations, computation, and non-asymptotic properties of these estimators.
Algorithm solves robust linear regression with block Lewis weights.
problem Group distributionally robust least squares problem.
method Algorithm based on geometric construction and block Lewis weights, using accelerated proximal methods.
result Improves over known methods for moderate accuracy regimes and matches state-of-the-art guarantees.
Efficiently estimates private least squares with linear error growth.
problem Private estimation of ordinary least squares with bounded residuals and leverage.
method Scaled noise added to a stable nonprivate estimator of the regression vector.
result Near-optimal accuracy guarantee with linear error growth in dimension.
Many problems in financial engineering involve the estimation of unknown conditional expectations across a time interval. Often Least Squares Monte Carlo techniques are used for the estimation. One method that can be combined with Least Squares Monte Carlo is the "Regress-Later" method. Unlike conventional methods wher…
Optimizes learning rates for kernel-based expectile regression.
problem Estimating conditional expectiles efficiently.
method Support vector machine type approach using Gaussian RBF kernels.
result Learning rates are minimax optimal with a logarithmic factor.
The paper improves least-squares regression learning rates for stronger norms.
problem Improving least-squares regression learning rates for stronger norms.
method Combining integral operator techniques with embedding properties.
result Learning rates for Sobolev norms without requiring the function to be in the hypothesis space.
New algorithm improves regression error bounds and accelerates performance for low noise.
problem Nonparametric least square regression in RKHS with optimal error bounds.
method Kernel Truncated Randomized Ridge Regression (KTRRR) with optimal generalization error bounds.
result Faster finite-time and asymptotic rates on low noise problems.
This work refutes the conventional wisdom and shows acceleration can be made robust for least squares regression.
problem The challenge of using fast gradient methods for stochastic optimization due to instability and error accumulation.
method Introduced an accelerated stochastic gradient method for least squares regression.
result Proves accelerated stochastic gradient descent achieves minimax optimal statistical risk faster than SGD.
Study optimal rates for spectral algorithms in Hilbert spaces.
problem Regression problems over separable Hilbert spaces with square loss.
method Investigate spectral/regularized algorithms including ridge, principal component, and gradient methods.
result Prove optimal, high-probability convergence results in terms of norms.
Least Squares Estimators are suboptimal for 5D convex functions.
problem Suboptimality of Least Squares Estimators in estimating multidimensional convex functions.
method Analysis of natural subclasses of convex functions in random and fixed design settings.
result Risk of LSE is n−2/d while minimax risk is n−4/(d+4) for d≥5. Optimal multiscale learning of linear operators
problem Statistical and computational limits of learning bounded linear operators between Sobolev spaces
method Reformulate as an infinite-dimensional matrix regression problem with heterogeneous multiscale structure
result Establish minimax rates and construct a finite-resolution blockwise least-squares estimator attaining these rates
Study on consistency of ML methods for moving objects in non-stationary environments.
problem Consistency of machine learning methods for moving objects in non-stationary environments.
method Least squares, ridge regression, and ℓs-penalized least squares methods under non-stationary spatial-temporal sampling. result Consistency and asymptotic normality of the estimates under weak conditions.
We prove statistical rates of convergence for kernel-based least squares regression from i.i.d. data using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is related to Kernel Partial Least Squares, a regression method that combines supervised dimensio…
Paper introduces ℓ-DER for regression tasks using morphological operators and convex-concave procedure.
problem Developing a universal approximator for regression tasks.
method Introduces ℓ-DER model, trains it using a convex-concave procedure (CCP) to minimize least-squares. result Outperforms other hybrid morphological models and state-of-the-art approaches.
Improved regression analysis using Padé approximants with new residuals and regularization.
problem Improving regression analysis with Padé approximants for accuracy and avoiding overfitting.
method New residuals in least squares method, system of linear equations for rational functions, Tikhonov regularization.
result Demonstrated efficiency in practical cases from physics and reliability theory.
Kernel methods with random projections improve least-squares regression efficiency.
problem Efficiently solving least-squares regression problems in high-dimensional spaces.
method Kernel conjugate gradient methods with randomized sketches and Nyström subsampling.
result Optimal generalization and computational advantages with proportional projection dimensions.
It is shown that the the popular least squares method of option pricing converges even under very general assumptions. This substantially increases the freedom of creating different implementations of the method, with varying levels of computational complexity and flexible approach to regression. It is also argued that…
Improved kernel ridge regression using conjugate gradients.
problem Efficiently solving large-scale kernel ridge regression problems.
method Structured Gaussian regression model with low-rank approximation and conjugate gradients.
result Enhanced approximation of kernel ridge regressor/Gaussian process posterior mean.
Least squares regression shows unexpected double descent in under-parameterized models.
problem Understanding the generalization of under-parameterized models in regression.
method Analyzing the spectrum and eigenvectors of the sample covariance matrix.
result Least squares regression can exhibit a peak in generalization in the under-parameterized regime, contrary to previous explanations.
Proposes a new regression method using Lp-norms for non-Gaussian noise.
problem Non-Gaussian noise in residuals affects the performance of local least squares regression.
method Introduces local polynomial Lp-norm regression, replacing weighted least squares with weighted Lp-norm estimation. result Demonstrates superior performance over local least squares in one-dimensional data and higher dimensions.
Improved SGD for non-strongly-convex regression with faster convergence.
problem Non-strongly-convex least squares regression problems.
method Modified accelerated gradient descent.
result Achieves optimal prediction error rates of O(d/t) and forgets initial conditions faster to O(d/t2). Develops a distributed least squares approximation method for regression problems.
problem Solving large-scale regression problems on distributed systems.
method Approximates local objective functions using a local quadratic form and combines estimators by weighted average.
result Statistically efficient combined estimator with one round of communication.
Study shows how mini-batch GD with random reshuffling affects least squares regression dynamics.
problem Analyzing the error dynamics of mini-batch GD with random reshuffling for least squares regression.
method Represented training and generalization errors through a sample cross-covariance matrix Z, compared with sample covariance matrix of original features X, and used linear scaling rule for analysis.
result Mini-batch GD with random reshuffling exhibits subtle step-size dependence not detectable by gradient flow analysis, converging to a limit dependent on the step size.