This paper explores how deep learning models can fit data exactly and why this is important.
problem Understanding why deep learning models can fit data exactly and generalize well.
method Interpolation and over-parameterization as key themes to understand deep learning.
result Interpolation and over-parameterization are crucial for deep learning models to fit data exactly and generalize well.
Study of graphs interpolating curve and pants graphs, providing formulae and geometry classifications.
problem Understanding the large-scale geometry of graphs connecting curve and pants graphs.
method Developed explicit formulae for quasi-flat ranks and classified geometries using twist-free graphs of multicurves.
result Explicit formulae for quasi-flat ranks and classification of geometries into hyperbolic, relatively hyperbolic, and thick cases.
Loose bounds found for least-norm interpolant in over-parameterized settings.
problem Failures of model-dependent generalization bounds for least-norm interpolation.
method Analysis of generalization performance of least-norm linear regressor in over-parameterized regime.
result Generalization bounds for least-norm interpolant can be very loose, even when true excess risk goes to zero.
New method interpolates high-dimensional scattered data using kernel theory.
problem Scattered data in high-dimensional spaces defy traditional distributional assumptions.
method Kernel interpolation framework based on integral operator theory.
result Spectra of kernel matrices predict performance of interpolation methods.
This paper introduces a new bound to explain generalization in over-parameterized models.
problem Understanding why some over-parameterized models generalize well while others do not.
method PAC-Chernoff bounds and smoothness measures based on large deviation theory.
result Interpolators with smoother structures generalize better, according to the new theoretical framework.
Study on biharmonic and interpolating sesqui-harmonic vector fields on para-Kähler--Norden manifolds.
problem Investigating higher-order harmonicity in pseudo-Riemannian geometry.
method Deriving first variations of bienergy and interpolating sesqui-energy functionals, characterizing biharmonic and interpolating sesqui-harmonic vector fields.
result Explicit characterizations and examples of vector fields satisfying biharmonic and interpolating sesqui-harmonic conditions.
Near-interpolating models grow norms quickly, affecting generalization.
problem Understanding the trade-off between interpolation and generalization in near-interpolating models.
method Random matrix theory and eigendecay analysis of data covariance matrix.
result Near-interpolating models exhibit rapid norm growth and worse generalization trade-offs.
The paper studies optimal transport in linear quadratic systems and derives interpolation inequalities.
problem Optimal transport problem in Linear Quadratic optimal control systems.
method Well-posedness of the Monge problem, regularity of optimal transport map, displacement interpolation of measures.
result Derivation of general interpolation inequalities for entropy functionals.
IIC provides a PAC-Bayes bound for interpolating models, revealing factors affecting generalization.
problem Theoretical challenges in understanding overparameterized models and their performance.
method PAC-Bayesian perspective applied to the Interpolating Information Criterion (IIC).
result Test error for overparameterized models achieving zero training error depends on various factors.
Geometric theory connects machine learning classifiers to differential geometry.
problem Classifying data points in machine learning.
method Mapping binary classification to vector bundles and differential geometry.
result Harmonic interpolation solves RKHS interpolation problems.
This paper analyzes the interpolation error of nonlinear Attention compared to linear regression.
problem Understanding the interpolation error of nonlinear Attention in high-dimensional settings.
method Derives explicit expressions for mean-squared interpolation error using signal-plus-noise model and random matrix theory.
result Nonlinear Attention generally incurs a larger interpolation error than linear regression, but this gap can be reversed with structured signals.
Paper develops SINNOs for approximating stochastic processes.
problem Approximating stochastic processes with neural networks.
method Developed stochastic interpolation neural network operators (SINNOs) with random coefficients.
result Established boundedness, interpolation accuracy, and approximation capabilities of SINNOs.
Paper investigates optimal interpolation methods in linear regression.
problem Understanding when interpolating methods generalize well in linear regression.
method Investigates optimal response-linear interpolators using functions linear in the response variable.
result Provides a closed-form expression for the optimal interpolator and shows it can be derived as the limit of gradient descent.
A new tradeoff between regularization and sharpness improves model performance in overparameterized settings.
problem Improving model performance in overparameterized settings with minimum-norm interpolators.
method Proposes a regularization-sharpness tradeoff for overparameterized linear regression with an ℓ^p penalty.
result Empirical validation shows the tradeoff terms can distinguish performant linear interpolators.
The paper analyzes the generalization error of min-norm interpolators in transfer learning with limited test samples.
problem Characterizing the generalization error of min-norm interpolators in transfer learning with limited test samples.
method Characterizes the bias and variance of pooled min-ℓ2-norm interpolation under covariate shift and model shift. result Shows that adding data can hurt when SNR is low and is beneficial at higher SNR levels under certain conditions.
Study on learning properties of scale-dependent kernels controlling stability and error.
problem Understanding the learning properties of scale-dependent kernels in nonparametric ridge-less least squares.
method Combines probabilistic results with interpolation theory to analyze stability and error.
result Different regimes of learning error depending on sample size and data dimension.
In this article, a proof of the interpolation inequality along geodesics in p-Wasserstein spaces is given. This interpolation inequality was the main ingredient to prove the Borel-Brascamp-Lieb inequality for general Riemannian and Finsler manifolds and led Lott-Villani and Sturm to define an abstract Ricci curvature…
Study investigates overparametrization in survival models, revealing complex loss behavior.
problem Understanding overparametrization in survival models through interpolation.
method Defined interpolation and finite-norm interpolation, rigorously analyzed four survival models.
result Overparametrization can lead to improved performance in survival models, contrary to classical learning theory.
New insights into Nadaraya-Watson interpolators show varied generalization behaviors.
problem Understanding generalization of interpolating predictors, especially in noisy data.
method Revisiting Nadaraya-Watson estimator with a single hyperparameter.
result Multiple overfitting behaviors exist, ranging from catastrophic to tempered.
Randomly sampled interpolators achieve zero generalization error with enough data.
problem Understanding the high generalization ability of machine learning models.
method Algebraic geometry tools to prove zero generalization error for random interpolators.
result Generalization error of randomly sampled interpolators becomes zero once the number of training samples exceeds a geometric threshold.
Lower bound proves ridgeless regression performs poorly near interpolation threshold.
problem Proving performance of ridgeless regression near interpolation threshold.
method Distribution-independent lower bound for mean squared error in noisy ridgeless linear regression.
result Lower bound implies ridgeless regression performs poorly near interpolation threshold.
Many modern machine learning models are trained to achieve zero or near-zero training error in order to obtain near-optimal (but non-zero) test error. This phenomenon of strong generalization performance for "overfitted" / interpolated classifiers appears to be ubiquitous in high-dimensional data, having been observed …
We analyze ridge interpolators in correlated factor regression models using RDT.
problem Performance analysis of ridge interpolators in correlated factor regression models.
method Utilizing Random Duality Theory (RDT), we obtain precise closed form characterizations of optimization problems.
result Ridge interpolators can smooth out the excess prediction risk and exhibit double-descent behavior.
Study optimizes linear regression analysis for high-dimensional settings.
problem Understanding high-dimensional linear regression with interpolation and regularization.
method Localized uniform convergence analysis of optimistic rates for linear regression.
result Recover guarantees for ridge and LASSO regression under random designs.
The study analyzes robustness of estimators in linear models with adversarial errors.
problem Analyzing robustness of estimators in linear models with adversarial errors.
method Develops a general theory for minimum norm interpolating estimators and RERM in linear models without conditions on errors.
result Quantitative bound for the prediction error relating it to Rademacher complexity, norm of minimum norm interpolator of errors, and subdifferential size.
Unified framework explains why overfitting is benign in interpolating learning.
problem Understanding why overfitting is benign in highly overparameterized models.
method Spectral-transport stability framework.
result Sharp benign-overfitting criterion and explicit phase-transition rates.
The spectral flow theorem is applied to operators on finite intervals.
problem Operators on finite intervals without boundary conditions are not Fredholm.
method Interpolation theory is used to define boundary conditions making the operators Fredholm. The spectral flow theorem is applied to find the Fredholm index.
result The Fredholm index is given by the spectral flow of the operator path.
New proof of timelike minimal surfaces using split-harmonic maps.
problem Interpolating a split-Fourier curve to a timelike minimal surface.
method Using split-harmonic maps to solve the singular Björling problem.
result Solved the interpolation problem for timelike minimal surfaces.
Deep neural networks excel at learning the training data, but often provide incorrect and confident predictions when evaluated on slightly different test examples. This includes distribution shifts, outliers, and adversarial examples. To address these issues, we propose Manifold Mixup, a simple regularizer that encoura…
Two-layer neural networks can overfit without increasing risk when data is noisy.
problem Understanding why neural networks can overfit without increasing risk in noisy data.
method Combining bias and variance analysis in a high-dimensional setting.
result The excess learning risk of the interpolator decays under mild conditions.
The monotonic linear interpolation in deep networks often leads to plateaus, revealing biases in optimization.
problem Plateaus in the optimization landscape of deep networks during monotonic linear interpolation.
method Investigated monotonic linear interpolation on deep neural networks, focusing on biases in weights and biases.
result Interpolating weights and biases differently can lead to significant differences in loss and accuracy, revealing biases in optimization.
Mixup improves model performance by interpolating random training examples.
problem Overfitting in machine learning models.
method Mixup is a regularization procedure that linearly interpolates random pairs of training examples.
result Mixup works well from a statistical learning theory perspective.
Adversarial robustness has become a central goal in deep learning, both in the theory and the practice. However, successful methods to improve the adversarial robustness (such as adversarial training) greatly hurt generalization performance on the unperturbed data. This could have a major impact on how the adversarial …
Interpolating models can have heavy-tailed risk, leading to rare but severe errors.
problem Interpolating models' tail risk is poorly understood, affecting rare but impactful errors.
method Large-deviation methods to study the fragility of high-dimensional linear interpolators.
result Ridgeless regression exhibits heavy-tailed risk, while ridge-regularized estimators have better tail behavior.
New method interpolates training data and is consistent for various data distributions.
problem Establishing generalization guarantees for ensemble methods in the interpolating regime.
method Developed manifold-Hilbert kernel for Riemannian manifolds and used it in ensemble classification.
result Consistent ensemble classification method for broad data distributions.
New sampling method uses stochastic interpolants and FBSDEs.
problem Sampling from high-dimensional distributions with unnormalized densities.
method Stochastic interpolants and FBSDEs to define and solve diffusion process.
result Effective sampling from challenging distributions.
New law explains why deep learning models often have more parameters than needed.
problem Why deep learning models often have more parameters than classical theory suggests.
method Proved a universal law of robustness for smooth interpolation.
result Smooth interpolation requires d times more parameters than mere interpolation.
NODEs with explicit time dependence can interpolate and generalize like piecewise-constant estimators.
problem Learning from finite datasets with neural ODEs.
method Control-theoretic perspective applied to semi-autonomous NODEs.
result SA-NODEs can interpolate and satisfy SCC, leading to generalization rates similar to histogram and nearest-neighbor estimators.
Following Kobayashi, we consider Griffiths negative complex Finsler bundles, naturally leading us to introduce Griffiths extremal Finsler metrics. As we point out, this notion is closely related to the theory of interpolation of norms, and is characterized by an equation of complex Monge--Ampère type, whose correspondi…
Ricci flow stabilizes hyperbolic 3-manifolds near the hyperbolic metric.
problem Stability of Ricci flow on hyperbolic 3-manifolds.
method Normalized Ricci-DeTurck flow with exponential convergence to the hyperbolic metric.
result Normalized Ricci-DeTurck flow converges exponentially to the hyperbolic metric.
One of the most well-known results in the theory of optimal transportation is the equivalence between the convexity of the entropy functional with respect to the Riemannian Wasserstein metric and the Ricci curvature lower bound of the underlying Riemannian manifold. There are also generalizations of this result to the …
New learning rates for embeddings in RKHSs, even when the target is not Hilbert-Schmidt.
problem Applying conditional mean embeddings to complex ML/RL settings with infinite-dimensional RKHSs.
method Developed novel learning rates using interpolation theory for RKHSs, derived explicit adaptive rates for sample estimator.
result Achieved uniform convergence rates in the output RKHS for certain parameter regimes.
Study reveals phase transition in neural networks near interpolation.
problem Understanding generalization and learning transitions in neural networks.
method Effective theory for approximating Bayes-optimal generalisation error.
result Unveils a discontinuous phase transition between universal and specialisation phases.
We prove a new generalization bound that shows for any class of linear predictors in Gaussian space, the Rademacher complexity of the class and the training error under any continuous loss ℓ can control the test error under all Moreau envelopes of the loss ℓ. We use our finite-sample bound to directly recover…
A common strategy to train deep neural networks (DNNs) is to use very large architectures and to train them until they (almost) achieve zero training error. Empirically observed good generalization performance on test data, even in the presence of lots of label noise, corroborate such a procedure. On the other hand, in…
Generalized complex geometry, introduced by Hitchin, encompasses complex and symplectic geometry as its extremal special cases. We explore the basic properties of this geometry, including its enhanced symmetry group, elliptic deformation theory, relation to Poisson geometry, and local structure theory. We also define a…
New principle controls graph-informed adversarial discrepancies.
problem Graph-informed adversarial learning for interpolative divergences.
method Proves infimal subadditivity for interpolative divergences.
result Graph-informed adversarial learning is justified for interpolative divergences.
New learning rates derived for Tikhonov-regularized problems without kernel assumptions.
problem Learning rates for Tikhonov-regularized learning problems.
method Minimax adaptive rates derived using Fourier isocapacitary condition and interpolation theory.
result Derivation of minimax adaptive rates without requiring kernel assumptions.