Optimizes Lasso hyperparameters using leave-one-out CV.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Optimizes hyperparameter tuning for models using approximate leave-one-out cross-validation.
ALO-CV approximates leave-one-out error in proportional regime.
The paper improves ALO for -regularized models.
The paper proves LOO CV is reliable under estimator stability.
A fast method for LOOCV in k-NN regression reduces computation time.
Improved LOO cross-validation for function approximation.
Paper accelerates conformal prediction by using approximate leave-one-out estimators.
Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but unfortunately, LOO does not scale well to large datasets. We propose a combination of u…
We show how to adjust the coefficient of determination () when used for measuring predictive accuracy via leave-one-out cross-validation.
Study improves understanding of non-differentiable penalties in high-dimensional settings.
Algorithm identifies and corrects noisy labels using Gaussian process regression.
A new method improves robustness and efficiency of Bayesian LOO-CV.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for stable predictors in the context of risk assessment. The notion of stability has been first introduced by \cite{DEWA79} and extended by \cite{KEA95}, \cite{BE01} and \cite{KUNIY02} to characterize cla…
New method optimizes hyperparameters for non-smooth problems efficiently.
New methods improve anomaly detection with reduced false positives.
The future predictive performance of a Bayesian model can be estimated using Bayesian cross-validation. In this article, we consider Gaussian latent variable models where the integration over the latent values is approximated using the Laplace method or expectation propagation (EP). We study the properties of several B…
Enhances polynomial chaos models with uncertainty intervals.
Proposes DeGLIF to denoise graph data for label noise robustness.
RandALO speeds up risk estimation for large datasets.
The study evaluates different parameter selection methods for Gaussian process interpolation.
We speed up Gaussian process cross-validation calculations and improve model diagnostics.
We consider the problem of estimating the parameters of the covariance function of a Gaussian process by cross-validation. We suggest using new cross-validation criteria derived from the literature of scoring rules. We also provide an efficient method for computing the gradient of a cross-validation criterion. To the b…
The paper improves confidence intervals for test error using cross-validation.
Risk estimation is at the core of many learning systems. The importance of this problem has motivated researchers to propose different schemes, such as cross validation, generalized cross validation, and Bootstrap. The theoretical properties of such estimates have been extensively studied in the low-dimensional setting…
Bayesian model averaging, model selection and its approximations such as BIC are generally statistically consistent, but sometimes achieve slower rates og convergence than other methods such as AIC and leave-one-out cross-validation. On the other hand, these other methods can br inconsistent. We identify the "catch-up …
Weighted SVM (or fuzzy SVM) is the most widely used SVM variant owning its effectiveness to the use of instance weights. Proper selection of the instance weights can lead to increased generalization performance. In this work, we extend the span error bound theory to weighted SVM and we introduce effective hyperparamete…
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for subagged estimators, both for classification and regressor. General loss functions and class of predictors with both finite and infinite VC-dimension are considered. We slightly generalize the formali…
Backward Conformal Prediction offers flexible control over prediction set sizes while ensuring coverage guarantees.
We study the problem of out-of-sample risk estimation in the high dimensional regime where both the sample size and number of features are large, and can be less than one. Extensive empirical evidence confirms the accuracy of leave-one-out cross validation (LO) for out-of-sample risk estimation. Yet, a un…
Cross-validation (CV) is a technique for evaluating the ability of statistical models/learning systems based on a given data set. Despite its wide applicability, the rather heavy computational cost can prevent its use as the system size grows. To resolve this difficulty in the case of Bayesian linear regression, we dev…
This paper introduces Kernel-based Information Criterion (KIC) for model selection in regression analysis. The novel kernel-based complexity measure in KIC efficiently computes the interdependency between parameters of the model using a variable-wise variance and yields selection of better, more robust regressors. Expe…
We propose a novel algorithm for greedy forward feature selection for regularized least-squares (RLS) regression and classification, also known as the least-squares support vector machine or ridge regression. The algorithm, which we call greedy RLS, starts from the empty feature set, and on each iteration adds the feat…
UDM reparameterization improves language model generation.
Decoding, ie prediction from brain images or signals, calls for empirical evaluation of its predictive power. Such evaluation is achieved via cross-validation, a method also used to tune decoders' hyper-parameters. This paper is a review on cross-validation procedures for decoding in neuroimaging. It includes a didacti…
Paper analyzes robust matrix completion with efficient nonconvex method and leave-one-out analysis.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for empirical risk minimizers. In the general setting, we prove sanity-check bounds in the spirit of \cite{KR99} \textquotedblleft\textit{bounds showing that the worst-case error of this estimate is not m…
FastMuyGPs speeds up GP predictions for large datasets.
Adaptive coverage policies improve conformal prediction accuracy.
Framework designs antiviral drugs using deep learning and RL.
New method improves online nonparametric estimators with minimal extra computation.
Bayesian EM method improves ridge regression tuning without LOOCV's limitations.
The present paper provides a new generic strategy leading to non-asymptotic theoretical guarantees on the Leave-one-Out procedure applied to a broad class of learning algorithms. This strategy relies on two main ingredients: the new notion of stability, and the strong use of moment inequalities. stability e…
Analysis of cross-validation for early-stopped gradient descent in high-dimensional regression.
LOOCV is often useful for analyzing small, structured experimental designs.
Consider the following class of learning schemes: where and denote the feature and response variable …
LOO-StabCP speeds up CP for multiple predictions.
We investigate the issue of model selection and the use of the nonconformity (strangeness) measure in batch learning. Using the nonconformity measure we propose a new training algorithm that helps avoid the need for Cross-Validation or Leave-One-Out model selection strategies. We provide a new generalisation error boun…