Two new algorithms recover ridge lines from point clouds with convergence guarantees.
problem Extracting filamentary structure from point clouds.
method Proposes two novel algorithms with convergence guarantees.
result The algorithms can asymptotically recover the full ridge set.
Study parallel and dual surfaces of cuspidal edges, defining ridge points and clarifying geometric relations.
problem Understanding the geometric properties of cuspidal edges and their duals.
method Analyzing principal curvature and direction, defining ridge points, and examining geometric relations.
result Clarified relations between singularities of parallel and dual surfaces and cuspidal edges.
We study the problem of estimating the ridges of a density function. Ridge estimation is an extension of mode finding and is useful for understanding the structure of a density. It can also be used to find hidden structure in point cloud data. We show that, under mild regularity conditions, the ridges of the kernel den…
OKRidge solves sparse ridge regression problems for nonlinear systems.
problem Identifying sparse governing equations for nonlinear dynamical systems.
method OKRidge algorithm using saddle point formulation and ADMM-based approach with efficient proximal operators.
result OKRidge achieves provable optimality with significantly faster run times than Gurobi.
New insights into ridge regression with correlated data, improving risk prediction.
problem Understanding and predicting risk in ridge regression with correlated samples.
method Random matrix theory and free probability for asymptotic analysis; modified GCV estimator (CorrGCV) for unbiased prediction.
result GCV estimator fails for out-of-sample risk with correlated data; CorrGCV provides an unbiased estimator.
ASkotch solves large-scale KRR faster and better than existing methods.
problem Challenges in scaling full Kernel Ridge Regression (KRR) to large datasets.
method ASkotch: A scalable, accelerated, iterative method for full KRR.
result ASkotch provides better solutions faster than state-of-the-art solvers for full and inducing points KRR.
Estimates modes and ridges in mixed Euclidean and directional spaces.
problem Estimating local modes and density ridges in product spaces combining Euclidean and directional metrics.
method Extends mean shift algorithm to product spaces, addressing challenges in generalization.
result Established convergence of the proposed methods and demonstrated effectiveness on real-world datasets.
Localized sketching improves matrix multiplication and ridge regression complexity.
problem Efficiently approximate matrix multiplication and ridge regression with limited data availability.
method Localized sketching matrices for block diagonal structure, reducing sample complexity.
result Localized sketching achieves sample complexity matching global sketching methods.
Deterministic algorithm estimates ridge regression with minimal space.
problem Estimating ridge regression solutions efficiently.
method Deterministic space-efficient algorithm using Frequent Directions.
result First o(d2) space deterministic streaming algorithm with guaranteed error. Two methods solve kernel ridge regression problems efficiently.
problem Solving kernel ridge regression problems with large datasets.
method RPCholesky and KRILL preconditioning techniques.
result Efficient solutions to KRR problems with strong guarantees.
Paper improves understanding of random Fourier features for kernel ridge regression.
problem Understanding statistical properties of random Fourier features for kernel ridge regression.
method Spectral matrix approximation approach to analyze random Fourier features.
result Proves statistical guarantees for kernel ridge regression using random Fourier features.
Method finds hidden structures in noisy data.
problem Finding unexpected or unknown structures in noisy point clouds.
method First find ridges in estimated density, then filter using Hessian eigenvalues.
result Outputs well-defined sets of dimensions lower than ambient space.
The classical Lusternik-Schnirelman-Borsuk theorem states that if a d-sphere is covered by d+1 closed sets, then at least one of the sets must contain a pair of antipodal points. In this paper, we prove a combinatorial version of this theorem for hypercubes. It is not hard to show that for any cover of the facets of a …
A new approach to Kernel Ridge Regression using partitioning.
problem Efficiently estimating KRR for large datasets.
method Divide and conquer approach via partitioning of input space.
result Achieves optimal minimax rates and reduces approximation error.
New method improves prediction accuracy by using future test points.
problem Improving prediction accuracy by leveraging future test points.
method Transductive inference techniques from semi-parametric inference.
result Non-asymptotic upper bounds on transductive prediction errors.
RF models implicitly regularize kernel methods as feature count increases.
problem Understanding implicit regularization in RF models.
method Random matrix theory applied to Gaussian RF models and KRR.
result The average RF predictor is close to a KRR predictor with an effective ridge.
Paper bounds the minimal rank for kernel ridge regression approximations.
problem Efficient memory and computation for kernel ridge regression.
method Lower bound on minimal rank for reliable prediction power.
result Nyström method's computational cost is almost linear in sample size.
Proposes landmark selection for kernel methods.
problem Selecting important landmarks from large training sets.
method Deterministic and randomized adaptive algorithm for landmark selection.
result Landmarks are related to the minima of kernelized Christoffel functions.
Analyzes merging vs. ensembling for multi-study prediction, showing transition point for better performance.
problem Choosing between merging or ensembling multiple studies for prediction.
method Analyzes ridge regression approaches, comparing merging and ensembling methods.
result There is a transition point where ensembling outperforms merging as cross-study heterogeneity increases.
STARK improves denoising of low-depth spatial transcriptomics images.
problem Denoising spatial transcriptomics images at ultra-low sequencing depths.
method Adaptive regularization with kernel ridge regression and graph Laplacian.
result STARK optimizes denoising performance over competing methods.
Two new ridge solutions improve BLS on added nodes, achieving better accuracy.
problem Improving the Broad Learning System (BLS) for new nodes.
method Proposed two ridge solutions for BLS output weights, updating efficiently.
result Proposed ridge solutions achieve better testing accuracy than original BLS.
Paper proposes a new landmark selection method for kernel ridge regression.
problem Efficient landmark selection for scalable kernel methods.
method Two-step approach: first computes importance scores, then clusters them into landmarks.
result Proposed method provides better accuracy and efficiency trade-offs.
A new algorithm for efficient kernel Nyström approximation.
problem Efficiently approximating large kernel matrices for machine learning.
method Recursive sampling of landmark points using ridge leverage scores.
result Scalable and accurate kernel approximation with linear runtime.
Ridge regularization simplifies model complexity in data science.
problem Overfitting in statistical models.
method Adding a penalty on the magnitude of coefficients.
result Effective in reducing model complexity and improving generalization.
Deterministic column sampling using ridge leverage scores provides accurate matrix sketches for ridge regression.
problem Regularizing ill-posed linear least-squares problems with small but non-zero coefficients.
method Deterministic column sampling using ridge leverage scores.
result Deterministic algorithm provides (1 + ε) error column subset selection and projection-cost preservation.
The paper analyzes a simple neural network model with algebraic methods.
problem Finding minima of a ridge-regularized mean squared error for ReLU perceptrons.
method Developed a Divide-Enumerate-Merge strategy using computational algebra.
result Identifies both isolated and connected minima of the RR-MSE.
Paper proposes FR algorithm to solve minimax optimization locally.
problem Gradient descent fails to find local minimax in minimax optimization.
method Follow-the-Ridge (FR) algorithm, addressing rotational behavior of gradient dynamics.
result FR algorithm provably converges to local minimax.
Improved ridge estimators avoid tuning parameters for high-dimensional data.
problem Difficulty in calibrating tuning parameters for ridge estimators.
method Developed modified ridge estimators that eliminate tuning parameters.
result Modified ridge estimators outperform standard methods in prediction accuracy.
Optimal ridge penalty can be negative or zero in high-dimensional data.
problem Overfitting in high-dimensional underdetermined linear regression.
method Simulations and real-life data analysis with minimum-norm estimator.
result Optimal ridge penalty can be negative, contradicting conventional wisdom.
New equivalences found between subsampling and ridge regularization methods.
problem Establishing precise structural and risk equivalences between subsampling and ridge regularization.
method Proved structural and risk equivalences between subsample ridge estimators and different ridge regularization levels and subsample aspect ratios.
result Optimally tuned ridge regression exhibits a monotonic prediction risk in the data aspect ratio.
Study ridge ensembles in proportional feature-to-sample size regime, proving risk equivalence and GCV consistency.
problem Characterizing and optimizing ridge ensembles in proportional feature-to-sample size regimes.
method Proportional asymptotics analysis, GCV for tuning, proving risk equivalence.
result Risk of optimal full ridgeless ensemble matches optimal ridge predictor's risk.
KIP meta-learning compresses datasets significantly.
problem Training data size and quality issues in machine learning.
method Kernel Inducing Points (KIP) for dataset compression.
result Significant reduction in dataset size with similar model performance.
We are concerned with an approximation problem for a symmetric positive semidefinite matrix due to motivation from a class of nonlinear machine learning methods. We discuss an approximation approach that we call {matrix ridge approximation}. In particular, we define the matrix ridge approximation as an incomplete matri…
Boosting ridge regression for high-dimensional data classification reduces computational cost and improves learning time.
problem High computational demand of inverting regularised covariance matrix in ridge regression for high-dimensional problems.
method Train an ensemble of ridge regressors in randomly projected subspaces, then combine them using adaptive boosting.
result Effective in terms of learning time and improved predictive performance in some cases.
Manifold learning has been successfully applied to a variety of medical imaging problems. Its use in real-time applications requires fast projection onto the low-dimensional space. To this end, out-of-sample extensions are applied by constructing an interpolation function that maps from the input space to the low-dimen…
Paper refines Mean Shift for stable density ridges, proving convergence to new geometric structure.
problem Theoretical mismatch between static and dynamic density ridges in SCMS.
method Introduces stable ridge, proves its convergence, and develops generalized SCMS framework.
result SCMS converges to stable ridge, providing a more accurate representation of data.
Short proof shows how ridge regression works with random data.
problem Understanding prediction error in ridge regression with random design.
method Combination of exchangeability arguments, matrix perturbation, and operator convexity.
result Elementary proof of prediction error without complex inequalities.
The paper examines how nonlinear transformations affect ridge sets in manifold learning.
problem Understanding the impact of nonlinear transformations on ridge sets in manifold learning.
method Examined the effects of nonlinear transformations on ridge sets using mathematical proofs and numerical experiments.
result The inclusion relationship $\cR(f\circ p)\subseteq \cR(p)$ holds for strictly increasing and concave transformations, and the Hausdorff distance between transformed and non-transformed ridge sets is smaller.
HARFE approximates sparse additive functions using random features and ridge regression.
problem Approximating high-dimensional sparse additive functions.
method Hard-ridge random feature expansion with sparse ridge regression and hard-thresholding pursuit.
result HARFE method converges with a given error bound and achieves lower error than other algorithms.
New insights into how neural networks learn features, especially when they are very wide.
problem Understanding how gradient flow in wide neural networks selects solutions, especially in the feature-learning regime.
method Axiomatizing the canonical regularizer as a function-space energy and lift, and deriving geodesic ridge for the feature-learning regime.
result Gradient flow in feature-learning networks biases towards ridge regularization, distorting the inductive bias and damaging pretrained networks.
Unified study of ridge regression structure, cross-validation, and acceleration.
problem Understanding and optimizing ridge regression in large-data settings.
method Unified large-data linear model analysis, cross-validation bias correction, sketching accuracy study.
result Unified understanding and improved methods for ridge regression.
New findings on how overfitting can be beneficial in ridge regression.
problem Understanding overfitting in overparameterized models.
method Extending previous results on linear regression to ridge regression, eliminating independence assumptions.
result Sharp bounds on the variance and bias terms, explaining optimal regularization in ridge regression.
MGD with early stopping tends to ridge regularization in least squares regression.
problem Characterizing the implicit regularization of MGD with early stopping.
method Continuous-time view of MGD (momentum gradient flow) and comparison with explicit ridge regularization.
result Under optimal tuning, the risk of MGF is no more than 1.54 times that of ridge.
Kernel balancing weights are generalized as KRRR, providing better confidence intervals for treatment effects.
problem Lack of generalization error, correct feature specification, and limited to average effects.
method Interpreting kernel balancing weights as KRRR, relaxing feature specification, and extending Gaussian approximation.
result KRRR provides strong generalization properties and justifies confidence sets for causal functions.
Kernel ridge regression imputation with consistent variance estimation for handling missing data.
problem Handling missing data in statistical analysis.
method Kernel ridge regression imputation combined with entropy method for variance estimation.
result Root-n consistency of the imputation estimator in a Sobolev space setting.
Regularized linear regression improves binary classification performance, especially with ridge and ℓ1 regularization.
problem Improving binary classification accuracy with noisy labels.
method Systematic study of regularization strengths on linear classifiers trained on noisy binary classification data.
result Ridge regression consistently improves classification error, while ℓ1 regularization can induce sparsity and ℓ∞ regularization can concentrate weights to two values. A new method for high-dimensional functional regression reduces multicollinearity and improves interpretability.
problem Multicollinearity, overfitting, and interpretability in high-dimensional functional linear models.
method Partition-based functional ridge regression framework.
result Improved numerical stability and enhanced interpretability without explicit variable selection.
Efficient variance estimation for kernel ridge regression.
problem Estimating variance in kernel ridge regression efficiently.
method Random projection approach to estimate variance.
result Optimal variance estimator for various kernels.