Batching stabilizes risk in high-dimensional linear regression models.
problem Stability and risk behavior in high-dimensional overparameterized linear regression.
method Minimum-norm overparameterized linear regression model with batch-partitioning.
result Optimal batch size is inversely proportional to noise level and overparametrization ratio, leading to stable risk behavior.
The study analyzes robustness of estimators in linear models with adversarial errors.
problem Analyzing robustness of estimators in linear models with adversarial errors.
method Develops a general theory for minimum norm interpolating estimators and RERM in linear models without conditions on errors.
result Quantitative bound for the prediction error relating it to Rademacher complexity, norm of minimum norm interpolator of errors, and subdifferential size.
Study shows unusual non-monotonic risk behavior in minimum-norm interpolants for various data scaling.
problem Understanding the risk behavior of minimum-norm interpolants in RKHS for different data scaling.
method Analysis of spectral properties of the random kernel matrix restricted to eigen-spaces of the population covariance operator.
result Minimum-norm interpolants in RKHS exhibit multiple descent in risk for d=nα with α∈(0,1). The paper analyzes the risk of a least squares estimator under a spike covariance model.
problem Risk analysis of the least squares estimator under a spike covariance model.
method Assumes spike covariance matrices, studies risk as d/nightarrow∞. result Risk of the minimum norm least squares estimator vanishes compared to the null estimator.
Estimates generalization error for two-layer ReLU NNs through minimum norm solutions.
problem Estimating generalization error for two-layer ReLU NNs trained by mean squared error.
method Uses minimum norm solutions and Neural Tangent Kernel (NTK) regime to derive generalization error bounds.
result Derives an a priori generalization error bound for two-layer ReLU NNs without requiring exponentially large number of neurons.
Minimum-norm solutions generalize well in over-parametrized neural networks.
problem Generalization error in over-parametrized neural networks.
method Analyzing three models: random feature model, two-layer neural network, and residual network.
result Generalization error for minimum-norm solutions is comparable to Monte Carlo rate, up to logarithmic terms.
Study shows how networks converge to minimum norm solutions with regularization.
problem Interpolating between known regions in shallow ReLU networks.
method Investigates empirical risk minimizers and weight decay regularizers.
result Empirical risk minimizers converge to minimum norm interpolants under specific conditions.
Optimal ridge penalty can be negative or zero in high-dimensional data.
problem Overfitting in high-dimensional underdetermined linear regression.
method Simulations and real-life data analysis with minimum-norm estimator.
result Optimal ridge penalty can be negative, contradicting conventional wisdom.
Exact expressions for double descent and implicit regularization in over-parameterized models.
problem Understanding the generalization error of over-parameterized models like deep neural networks.
method Surrogate random design to replace standard i.i.d. design, leading to exact expressions for mean squared error and implicit regularization.
result Exact non-asymptotic expressions for double descent and implicit regularization in over-parameterized models.
New method predicts brain activity using past states for more accurate source estimation.
problem Independent source estimates ignore temporal context of neuronal activity.
method Combines LSTM networks with Minimum-Norm Estimates (MNE) for context-dependent brain activity prediction.
result CMNE leads to more accurate source estimation compared to independent MNE.
Interpolation with Laplace kernel fails in low dimensions but succeeds in high dimensions.
problem Consistency of interpolation with Laplace kernels in low-dimensional settings.
method Minimum-norm interpolation in Reproducing Kernel Hilbert Space (RKHS) with Laplace kernel.
result Consistency of interpolation is a high-dimensional phenomenon.
Inflating the minimum norm interpolator improves linear regression generalization error.
problem Highly anisotropic covariances and diverging d/n in linear regression. method Inflating the minimum ℓ2 norm interpolator by a constant greater than one. result Inflating the minimum norm interpolator improves generalization error.
Statistical analysis of regularization in continual learning tasks.
problem Understanding how regularization affects model performance in sequential learning.
method Derivation of convergence rates, iterative update formula, and optimal hyperparameters for generalized ℓ2-regularization.
result Optimal hyperparameters balance forward and backward knowledge transfer, improving model performance.
New algorithm tackles high-dimensional contextual bandits without sparsity.
problem High-dimensional linear contextual bandit problem with large feature space.
method Proposes explore-then-commit (EtC) and adaptive explore-then-commit (AEtC) algorithms.
result Derives optimal rate for ETC algorithm and shows adaptive AEtC achieves it.
The hybrid Monte Carlo algorithm (HMCA) is applied for Bayesian parameter estimation of the realized stochastic volatility (RSV) model. Using the 2nd order minimum norm integrator (2MNI) for the molecular dynamics (MD) simulation in the HMCA, we find that the 2MNI is more efficient than the conventional leapfrog integr…
Estimates long-term effects using past experiments as instruments with many weak instruments.
problem Estimating long-term causal effects with limited short-term outcomes and many weak instruments.
method Nonparametric instrumental variable inference with many weak instruments, using past experiments as instruments.
result Automatic debiased machine learning estimators for linear functionals of the structural function and its minimum-norm projection are efficient in the many-weak-instruments regime.
Study shows interpolating predictor's risk is optimal in low-dimensional factor regression models.
problem Understanding the risk of interpolating predictors in high-dimensional factor regression models.
method Detailed finite-sample analysis of minimum-norm interpolating predictor's risk in factor regression models.
result The risk of the minimum-norm interpolating predictor approaches optimal benchmarks in low-dimensional factor regression models.
Task shift from classification to regression is possible in overparameterized linear models with limited additional data.
problem Transferability of latent knowledge from classification to regression in overparameterized linear models.
method Investigation of task shift in overparameterized linear regression, zero-shot and few-shot cases, with a focus on minimum-norm interpolation.
result Minimum-norm interpolators can transfer latent knowledge from classification to regression with limited additional data.
We introduce a norm on the space of test configurations, which we call the minimum norm. We conjecture that uniform K-stability with respect to this norm is equivalent to the existence of a constant scalar curvature Kähler metric. This notion of uniform K-stability is analogous to coercivity of the Mabuchi functional. …
The paper characterizes functions of shallow ReLU NN denoisers under minimal norm constraints.
problem Understanding the theoretical success of neural network denoisers.
method Characterization of functions realized by shallow ReLU NN denoisers under minimal norm constraints.
result The functions realized by shallow ReLU NN denoisers are contractive toward clean data points and generalize better than the empirical MMSE estimator at low noise levels.
Develops exact convex optimization formulations for neural networks.
problem Training two-layer neural networks with rectified linear units.
method Uses semi-infinite duality and minimum norm regularization to develop exact convex optimization formulations.
result Shows equivalence of ReLU networks trained with weight decay to block ℓ1 penalized convex models. The paper explores how over-parameterized linear regression models generalize without violating learning theory principles.
problem Understanding how over-parameterized linear regression models generalize without violating learning theory principles.
method The paper uses the predictive normalized maximum likelihood (pNML) learner to investigate the minimum norm solution of over-parameterized linear regression models.
result The model generalizes well when the test sample lies in a subspace spanned by eigenvectors associated with large eigenvalues of the training data.
Uniform convergence of interpolators proven for Gaussian data.
problem Interpolation learning in high-dimensional linear regression with Gaussian data.
method Generic uniform convergence guarantee in terms of Gaussian width.
result Consistency of interpolators for minimum-norm and near-minimal-norm cases.
The paper explores why a specific type of predictor works well in noisy data.
problem Understanding why a specific type of predictor (minimum-norm interpolator) works well in noisy data.
method The paper uses uniform convergence and zero-error predictors in a norm ball to explain the success of the minimum-norm interpolator.
result The minimum-norm interpolator is consistent, and this can be explained by uniform convergence of zero-error predictors in a norm ball.
Study shows double descent curve in high-dimensional linear regression with random projections.
problem Understanding the generalization performance in high-dimensional settings with random projections.
method Fixed prediction problem, ridge regression estimator, minimum norm least-squares fit, random matrix theory, asymptotic equivalents.
result Exhibit a double descent curve for high-dimensional linear regression with random projections.
Paper explains neural collapse in neural networks using a new model.
problem Understanding neural collapse in neural networks during training.
method Introducing the unconstrained layer-peeled model (ULPM) to prove gradient flow convergence to critical points of a minimum-norm separation problem.
result Proves that all critical points are strict saddle points except the global minimizers exhibiting neural collapse.
Study precise estimators for correlated data using RDT.
problem Analyzing estimators in correlated linear regression models.
method Utilized Random Duality Theory to characterize prediction risk.
result Precise closed form characterizations of estimators' risk.
Deep ResNets favor low bottleneck rank with proper hyperparameters.
problem Understanding the inductive bias of deep neural networks.
method Computed minimum-norm weights of a deep linear ResNet.
result Deep nonlinear ResNets have an inductive bias towards minimizing bottleneck rank.
Weight normalization and reparametrized gradient descent adaptively regularize weights and converge to minimum l2 norm solutions.
problem Adapting to non-convex weight normalization for convergence to minimum l2 norm solutions.
method Weight normalization and reparametrized projected gradient descent (rPGD) for overparametrized least-squares regression.
result rPGD converges close to the minimum l2 norm solution, even for far-from-zero initializations.
The paper shows how multi-task learning in neural networks is similar to kernel regression and Hilbert spaces.
problem Understanding the solutions to multi-task shallow ReLU neural network learning problems.
method Analyzing the properties of solutions to multi-task shallow ReLU neural network learning problems, proving uniqueness and equivalence to minimum-norm interpolation problems in Hilbert spaces.
result The solutions to multi-task neural network interpolation problems are almost always unique and coincide with the solution to a minimum-norm interpolation problem in a Sobolev (Reproducing Kernel) Hilbert Space.
Linear regression can overfit without harm when data is high-dimensional.
problem Understanding when a perfect fit to noisy training data in linear regression leads to accurate predictions.
method Characterization of linear regression problems with a minimum norm interpolating prediction rule that has near-optimal prediction accuracy.
result Overparameterization is essential for benign overfitting in this setting, and the number of unimportant prediction directions must exceed the sample size.
The paper analyzes early stopping in linear regression and shows it's equivalent to ridge regularization.
problem Understanding the effect of early stopping on linear regression models.
method Characterization of gradient descent dynamics and analysis of excess risk.
result Early stopped solution is equivalent to minimum norm solution for a generalized ridge regularized problem.
The study examines how shallow neural nets converge to training samples or manifold points during diffusion.
problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal ℓ2 norm, comparing score flow and diffusion flow. result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.
Deep networks generalize well without explicit regularization.
problem The apparent absence of overfitting in deep neural networks.
method Analysis of gradient descent dynamics and loss function minimization.
result Deep networks converge to minimum norm solutions, leading to good generalization.
This paper explores adaptive methods in over-parameterized linear regression.
problem Understanding why neural networks generalize well in over-parameterized settings.
method Characterizes two sub-classes of adaptive methods and their generalization performance.
result Adaptive methods in over-parameterized linear regression converge to the minimum norm solution.
Adversarial training improves linear regression solutions, offering robustness against small perturbations.
problem Vulnerability of linear models to adversarial perturbations.
method Formulated as a min-max problem, adversarial training minimizes the best solution under worst-case attacks.
result Adversarial training yields the minimum-norm interpolating solution in overparameterized models, equivalent to parameter shrinking methods in underparameterized models.
New method stabilizes machine learning for physics-informed inverse problems.
problem Reconstructing physical quantities from PDE-compliant measurements.
method Physics-informed learning with smooth inductive bias.
result PDE operators stabilize variance and prevent overfitting in fixed dimensions.
NTK neural networks are robust to adversarial attacks in nonparametric regression.
problem Adversarial robustness of neural networks in nonparametric regression.
method Gradient flow with early stopping for NTK neural networks, proving robustness in Sobolev spaces.
result NTK neural networks achieve optimal adversarial robustness rates in Sobolev spaces.
Kernel interpolation is inconsistent for norms with smoothness above a constant.
problem Inconsistency of kernel interpolation in reproducing kernel Hilbert spaces.
method Lower bounds for generalization error in Sobolev norms.
result Kernel interpolation is always inconsistent for norms with smoothness above a constant.
New algorithm finds minimum weight norm solutions in deep neural networks.
problem Training over-parameterized deep neural networks efficiently and improving generalization.
method Minnorm training method that minimizes the sum of weights' norms while fitting training data.
result Faster convergence to minimum-norm solutions and better generalization performance.
The conjugate gradient (CG) method is an efficient iterative method for solving large-scale strongly convex quadratic programming (QP). In this paper we propose some generalized CG (GCG) methods for solving the ℓ1-regularized (possibly not strongly) convex QP that terminate at an optimal solution in a finite numb…
Study tightens bounds for interpolating noisy data using minimum l1-norm.
problem Predicting noisy data with minimum l1-norm interpolation.
method Provided matching upper and lower bounds for prediction error.
result Tight consistency up to negligible terms for d≫n. Kernel method improves instrumental variable regression rates.
problem Nonparametric instrumental variable regression with weak instruments.
method Kernel-based two-stage least-squares method, strong L2 convergence analysis. result Minimax optimal rates for instrumental regression under standard assumptions.
Two active learning algorithms improve HSI classification using Fermat distances and harmonic label propagation.
problem Semi-supervised hyperspectral image classification with limited labeled data.
method Combines Fermat distances with Poisson-reweighted harmonic label propagation for active point selection.
result FALL and A-FALL algorithms enhance labeling accuracy and scalability for large HSI scenes.
Study ridge regression for non-identically distributed data with varying variances.
problem Investigate high-dimensional regression with non-identical data variance.
method Propose a random effect model and use tools from random matrix theory.
result Highlight the double descent phenomenon in high-dimensional regression for certain variance profiles.
Study shows minimizing the norm of the ERM solution stabilizes kernel ridge-less regression.
problem Stability of kernel ridge-less regression.
method Minimizing the norm of the ERM solution to minimize CV stability.
result Interpolating solution with minimum norm minimizes CV stability.
Kernel Ridgeless Regression can generalize well without explicit regularization.
problem The challenge of achieving generalization in Kernel Ridgeless Regression without additional regularization.
method Minimum-norm interpolated solutions with high-dimensional data, curvature of kernel function, and favorable data geometry.
result Implicit regularization leads to good generalization in Kernel Ridgeless Regression.
A new tradeoff between regularization and sharpness improves model performance in overparameterized settings.
problem Improving model performance in overparameterized settings with minimum-norm interpolators.
method Proposes a regularization-sharpness tradeoff for overparameterized linear regression with an ℓ^p penalty.
result Empirical validation shows the tradeoff terms can distinguish performant linear interpolators.