Regularization effect found in neural feature alignment.
problem Implicit regularization in deep learning models.
method Geometrical viewpoint and analysis of Rademacher complexity.
result Neural features align along task-relevant directions, leading to regularization.
We give asymptotically tight estimates of tangent space variation on Riemannian submanifolds of Euclidean space with respect to the local feature size of the submanifolds. We show that the result follows directly from structural properties of local feature size of the Riemannian submanifold and some elementary Euclidea…
Efficiently applies NTK to large-scale datasets using random features.
problem Computational limitations of kernel methods for large-scale datasets.
method Proposes a sketching-based algorithm combining random features of arc-cosine kernels to construct an efficient feature map of the NTK.
result Achieves comparable error bounds to exact kernel methods but with significantly reduced feature dimensionality.
Gradient descent learns useful features even in the NTK regime.
problem The ability of neural networks to learn useful features.
method Local convergence analysis of gradient descent with regularization.
result Gradient descent can capture ground-truth directions for feature learning even after the loss threshold is reached.
Regularizes 3D inverse scattering with tangent-point energy for better solutions.
problem Ill-conditioned inverse obstacle scattering problems in 3D.
method Tikhonov regularization using tangent-point energy to penalize surface roughness and ensure well-posedness.
result Regularized solutions converge to true solution as noise level decreases.
New insights into model robustness for random features and NTK models.
problem Understanding and distinguishing robustness in machine learning models.
method Analyzing empirical risk minimization in random features and NTK models.
result Random features models are not robust under any degree of over-parameterization, even when satisfying the universal law of robustness.
The paper analyzes how deep models memorize spurious features.
problem Understanding how deep models memorize spurious features in training data.
method Characterizes spurious feature memorization via model stability and feature alignment.
result Memorization of spurious features weakens as generalization capability increases.
Random features improve neural operators' generalization properties.
problem Improving generalization of neural operators.
method Unified framework for spectral regularization techniques and operator-valued kernels.
result Established optimal learning rates and required number of neurons.
Deep neural networks show some layers better align with data than others.
problem Understanding why some layers in deep neural networks better align with data.
method Introducing the Equilibrium Hypothesis to connect alignment pattern to signal propagation.
result The Equilibrium Hypothesis explains the ascent-descent pattern of alignment in deep neural networks.
New method finds sparse networks without labels, improving performance.
problem Sparse connectivity in neural networks to reduce memory and energy demands.
method Neural Tangent Transfer method to find sparse networks without labels.
result Sparse networks achieve higher classification performance and faster convergence.
Model shows feature learning can improve neural scaling laws for hard tasks.
problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.
New PINN architectures learn high-frequency features using Fourier features.
problem PINNs struggle with high-frequency or multi-scale features.
method Employ spatio-temporal and multi-scale random Fourier features.
result Effective PINN models for multi-scale PDEs.
Gradient descent aligns neural feature matrices with pre-activation tangent features.
problem Understanding neural feature learning mechanisms.
method Analytical proof of alignment between weight matrices and pre-activation tangent features.
result Derivative alignment occurs almost surely in high-dimensional settings.
TCT improves federated learning by convexifying neural networks.
problem Federated learning performance disparity due to nonconvexity.
method Train-Convexify-Train procedure using neural tangent kernels.
result Up to +36% accuracy improvement on FMNIST and +37% on CIFAR10.
Proposes graph neural network layers for manifold-valued graphs.
problem Graphs with features in a Riemannian manifold.
method Diffusion layer and tangent multilayer perceptron.
result Outperforms state-of-the-art networks on Alzheimer's classification.
TKIL improves class-balanced performance in incremental learning.
problem Catastrophic forgetting in sequential learning tasks.
method Introduces Tangent Kernel for Incremental Learning (TKIL) based on Neural Tangent Kernel (NTK).
result TKIL achieves better overall accuracy and variance across classes.
Empirical study shows standard CNNs deviate from NTK predictions.
problem Understanding how standard finite-width CNNs behave compared to their infinite-width NTK counterparts.
method Empirical analysis of AlexNet and LeNet architectures.
result Standard CNNs deviate significantly from their NTK counterparts, but deviation decreases with wider networks.
This paper tightens bounds on the smallest eigenvalue of NTK for deep ReLU networks.
problem Analyzing the smallest eigenvalue of Neural Tangent Kernel for deep ReLU networks.
method Analyzing various quantities of independent interest, including lower bounds on the smallest singular value of hidden feature matrices and upper bounds on the Lipschitz constant of input-output feature maps.
result Tight bounds on the smallest eigenvalue of NTK matrices for deep ReLU nets, both in the limiting case of infinite widths and for finite widths.
Study on how adversarial training affects neural network kernels and robustness.
problem Understanding and improving adversarial robustness in neural networks.
method Empirical study of the evolution of the empirical Neural Tangent Kernel (NTK) under standard and adversarial training.
result Adversarial training leads to a new kernel that provides robustness, even when non-robust training is performed on top of it.
New kernels from neural networks show better performance than traditional methods.
problem Improving neural network performance on small datasets.
method Developed algebraic operations to create compositional kernels from neural network architectures.
result Compositional kernels achieve higher accuracy than neural tangent kernels and neural networks on small datasets.
Analyzes neural networks using linear models to understand their behavior.
problem Understanding multi-layer neural networks through linear models.
method Recalls and reviews four models: linear regression with concentrated features, kernel ridge regression, random feature model, and neural tangent model.
result Highlights limitations of linear theory and discusses approaches to overcome them.
We study the training and generalization of deep neural networks (DNNs) in the over-parameterized regime, where the network width (i.e., number of hidden nodes per layer) is much larger than the number of training data points. We show that, the expected 0-1 loss of a wide enough ReLU network trained with stochastic…
Wide neural networks with asymmetrical node scaling converge globally and learn features.
problem Global convergence and feature learning in over-parameterised shallow networks.
method Gradient-based optimisation of wide, shallow neural networks with asymmetrical node scaling.
result Gradient flow and gradient descent converge to a global minimum and learn features, unlike in the NTK parameterisation.
New insights into how overfitting affects neural networks' performance.
problem Understanding the generalization of overfitted two-layer neural networks.
method Analyzing the NTK model with ReLU activation, focusing on min ℓ2-norm solutions. result Generalization error of overfitted NTK models approaches a small limiting value, even with infinite neurons and samples.
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
problem Accelerating kernel methods and deep learning training.
method Generalizes leverage score sampling to neural networks and proves equivalence to neural tangent kernel ridge regression.
result Equivalence between regularized neural network and neural tangent kernel ridge regression under leverage score sampling initialization.
Theory explains why neural nets better learn Calabi-Yau metrics.
problem Learning Calabi-Yau metrics with neural networks.
method Developed a theory of metric flows in neural network space.
result Finite-width neural networks learn Calabi-Yau metrics better than fixed kernel methods.
Motivated by the study of collapsing Calabi-Yau threefolds with a Lefschetz K3 fibration, we construct a complete Calabi-Yau metric on C3 with maximal volume growth, which in the appropriate scale is expected to model the collapsing metric near the nodal point. This new Calabi-Yau metric has singular tangen…
Adaptive Group Lasso selects important features in neural networks.
problem Lack of interpretability in neural networks.
method Adaptive Group Lasso for feature selection.
result Consistent feature selection for neural networks with theoretical guarantee.
Poor approximators found in neural networks and random feature models.
problem Understanding why certain neural networks and models perform poorly in approximating functions.
method Established a scale separation of Kolmogorov width type and applied it to neural networks and random feature models.
result Reproducing kernel Hilbert spaces and two-layer neural networks are poor L2-approximators for certain functions. New random feature maps for Laplacian and related kernels.
problem Challenges in approximating the Laplacian kernel and its generalizations.
method Developed random feature maps for Laplacian and related kernels, providing efficient sampling schemes.
result Demonstrated the efficacy of these random feature maps on real datasets.
This paper introduces tangent display maps to simplify tangent category theory.
problem The category of smooth manifolds does not admit all pullbacks, complicating tangent category theory.
method Develops tangent display maps as a special class of maps well-behaved with respect to pullbacks.
result Tangent display maps simplify previous work in tangent categories and provide a new way to define open subobjects.
Paper constructs infinitely many tangent functors on diffeological spaces.
problem Tangent spaces in diffeological spaces are not uniquely defined.
method Introduced and constructed infinitely many non-isomorphic tangent functors.
result The choice of tangent functor is not unique outside smooth manifolds.
Deep Gaussian Processes are reinterpreted as deep trigonometric networks for tractable inference.
problem Challenging inference in DGPs due to intractable marginalization in latent function space.
method Viewing DGPs as deep trigonometric networks with Bochner's theorem, and using the wide limit with a bottleneck to translate DGPs into deep trigonometric networks.
result The weight space view yields the same effective covariance functions as obtained in function space, and varying prior distributions over network parameters is equivalent to employing different kernels.
FFN addresses spectral bias in neural value approximation, improving reinforcement learning performance.
problem Spectral bias in neural value approximation, leading to slow convergence and poor performance.
method Proposes Fourier feature networks (FFN) to overcome spectral bias by using a composite neural tangent kernel.
result FFN achieves state-of-the-art performance on challenging continuous control domains with faster convergence and better stability.
Neural networks can learn kernel machines with a data-dependent kernel.
problem Can neural networks in the rich feature learning regime learn a kernel machine?
method Demonstrated silent alignment effect in neural networks, showing they can learn a kernel machine with a data-dependent kernel.
result Neural networks in the rich feature learning regime can learn a kernel machine with a data-dependent kernel due to silent alignment.
BSA reduces network data by interpreting feature subspaces.
problem Interpreting feature subspaces of unlabeled network data.
method Barycentric Subspace Analysis (BSA) for unlabeled networks.
result BSA provides a more interpretable approach compared to PCA.
Study on triviality of tangent and generalized tangent bundles of manifolds.
problem Triviality of tangent and generalized tangent bundles of manifolds.
method Analyzing relations between tangent bundle TM and generalized tangent bundle TM=TM⊕T∗M of manifolds. result The generalized tangent bundle of a parallelizable manifold is trivial, but the converse is not always true.
While graph kernels (GKs) are easy to train and enjoy provable theoretical guarantees, their practical performances are limited by their expressive power, as the kernel function often depends on hand-crafted combinatorial features of graphs. Compared to graph kernels, graph neural networks (GNNs) usually achieve better…
This paper improves neural network generalization by dynamically learning kernel parameters.
problem Improving neural network generalization and adaptability.
method Diagonal adaptive kernel model that learns kernel eigenvalues and output coefficients during training.
result The diagonal adaptive kernel model significantly improves generalization over fixed-kernel methods.
Extends random feature analysis to spectral methods and improves learning rates.
problem Improving generalization properties of spectral methods in large-scale learning.
method Extends random feature analysis to a broad class of spectral regularization techniques, including gradient descent and Nesterov method.
result Obtains optimal learning rates for regularity classes, including those not in the RKHS.
Sprays on Frechet manifolds connect connections and tangent structures.
problem Characterizing linear symmetric connections on Frechet manifolds.
method Constructing connection maps and linear symmetric connections on tangent and second-order tangent bundles using sprays.
result A bijective correspondence exists between linear symmetric connections on tangent bundles and sprays.
The field of multiple view geometry has seen tremendous progress in reconstruction and calibration due to methods for extracting reliable point features and key developments in projective geometry. Point features, however, are not available in certain applications and result in unstructured point cloud reconstructions.…
The study proves sub-Riemannian manifolds cannot satisfy CD conditions unless they are Riemannian.
problem Characterizing sub-Riemannian manifolds that satisfy CD conditions. method Analysis of tangent cones and geodesics, construction of new RCD structures. result Sub-Riemannian manifolds are never CD(K,N) unless they are Riemannian. The paper proves Γ-convergence of discrete tangent-point energies to continuous energies and ropelength, with applications to biarc curves.
problem Proving convergence of discrete tangent-point energies to continuous energies and ropelength.
method Using biarc curves and interpolation, the paper proves Γ-convergence of discretized tangent-point energies to the continuous tangent-point energies and ropelength functional. result Discrete almost minimizing biarc curves converge to ropelength minimizers and minimizers of continuous tangent-point energies.
We propose a special deformation of the Sasaki metric on tangent and unit tangent bundle of a Hermitian locally symmetric manifold. Geodesics of this deformed metric have different projections on a base manifold for tangent or unit tangent bundle cases in contrast to usual Sasaki metric. Nevertheless, the projections o…
Neural networks generalize well despite overfitting due to high capacity.
problem Understanding why deep neural networks generalize well in overparameterized settings.
method High-dimensional asymptotic analysis of generalization under kernel regression with Neural Tangent Kernel.
result Test error exhibits non-monotonic behavior and can have additional peaks and descents in the overparameterized regime.
Study of normal and tangent maps to frontals.
problem Understanding geometric and dynamical properties of frontals.
method Geometrical and dynamical analysis of normal and tangent maps.
result Parallels of the tangent map to a frontal curve are right equivalent to the tangent map of a frontal curve under certain conditions.
Generalizes Connes's tangent groupoid for sub-Riemannian geometry.
problem Calculating tangent cones in sub-Riemannian geometry.
method Constructs a completion of MimesMimesR+imes using sub-Riemannian metric. result Calculates all tangent cones in Gromov-Hausdorff distance.