The paper studies ReLU layers, introducing new tools to understand their singular values and Gaussian mean width.
problem Understanding the role of ReLU layers in neural networks and their impact on network performance.
method Introducing ReLU singular values and Gaussian mean width of operators to study ReLU layers.
result ReLU singular values and Gaussian mean width provide metrics for distinguishing correctly and incorrectly classified data.
This paper shows how infinitely wide Tensor Networks converge to Gaussian Processes.
problem Understanding the relationship between Tensor Networks and Gaussian Processes.
method Analyzing the infinite-width limit of Tensor Networks and comparing them to Gaussian Processes.
result Infinitely wide Tensor Networks converge to Gaussian Processes, proving their equivalence.
Finite-width neural networks use non-Gaussian priors, extending Gaussian process theory.
problem Understanding the behavior of neural networks with finite width.
method Perturbative extension of Gaussian process theory to finite-width neural networks, tracking preactivation distributions.
result Non-Gaussian processes as priors in finite-width neural networks.
Wide neural networks can degrade performance, contrary to conventional wisdom.
problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.
Wide deep neural networks with Gaussian weights approximate Gaussian processes closely.
problem Understanding the approximation of deep neural networks with Gaussian weights to Gaussian processes.
method Established novel rates for the Gaussian approximation of random deep neural networks with Gaussian parameters and Lipschitz activation functions in the wide limit.
result The distance between the network output and the Gaussian approximation scales inversely with the width of the network.
ResNets approximate log-Gaussian at initialization, improving network performance.
problem Understanding the initialization behavior of deep neural networks like ResNets.
method Analyzing ReLU ResNets in the infinite-depth-and-width limit, showing log-Gaussian behavior.
result ResNets at initialization exhibit hypoactivation and interlayer correlations, which are not captured by Gaussian limits.
Unified theory for deep and recurrent networks using Gaussian processes.
problem Understanding capabilities and limitations of different network architectures.
method Unified derivation of mean-field theory from statistical physics of disordered systems.
result Gaussian processes yield identical Gaussian kernels for both architectures at a single time point or layer.
Study Gaussian-process limits of neural networks using tensor programs.
problem Understanding the behavior of neural networks as they approach infinite width.
method Quantitative analysis through tensor programs and Wasserstein distance.
result Explicit finite-width error bounds, showing convergence to Gaussian-process limits.
Study shows neural networks trained with GD converge to Gaussian processes with polynomial decay.
problem Understanding convergence of neural networks to Gaussian processes during training.
method Explicit upper bounds on quadratic Wasserstein distance between trained networks and Gaussian approximations.
result Polynomial decay of approximation error with network width and training time.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Paper analyzes infinite-width attention layers using Tensor Programs.
problem Capturing the infinite-width limit of attention layers.
method Tensor Programs framework to rigorously identify the limit distribution.
result Derives exact form of infinite-width limit distribution without Gaussian approximations.
Deep neural networks' infinite-width behavior approximated by Gaussian models.
problem Understanding the behavior of deep neural networks in the limit of infinite width.
method Using the Lindeberg exchange principle to approximate weights by Gaussian random variables.
result Quantitative bounds on the 2-Wasserstein distance between deep neural networks and Gaussian limits.
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.
Study of deep linear neural networks with proportional width and depth.
problem Lack of descriptive power in Gaussian limit of deep linear neural networks.
method Proportional infinite-width infinite-depth limit for deep linear neural networks.
result Characterization of limiting distribution as a nontrivial mixture of Gaussians.
New priors for deep neural networks converge to Gaussian processes.
problem Improving the performance and stability of deep neural networks.
method Extending prior distributions to include non-zero means and partially exchangeable priors, leading to a new Gaussian process model.
result The new Gaussian process model avoids pathologies and improves performance on regression problems.
Deep neural networks converge to Gaussian mixtures as layer width increases.
problem Understanding the distribution of outputs from deep neural networks.
method Proof and experiments with a simple model showing the convergence of neural network outputs to Gaussian mixtures.
result Neural networks converge to Gaussian mixtures as the width of the last hidden layer increases.
Develops EFT for ResNets, revealing limitations of kernel-only approach.
problem Limitations of kernel-only approach in deep neural networks.
method Collective kernel EFT for pre-activation ResNets based on G-only closure hierarchy. result Numerical findings show V4 equation residual accumulates to an O(1) error. Quantitative CLTs show neural network distributions converge to Gaussian as width increases.
problem Understanding the distribution of fully connected neural networks with random weights and biases.
method Analyzing the distribution of a fully connected neural network with random Gaussian weights and biases, proving quantitative bounds on normal approximations.
result The distance between a random fully connected network and the corresponding infinite width Gaussian process scales like n−γ for γ>0. It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width. This correspondence enables exact Bayesian inference for infinite width neural networks on regression tasks by means of eva…
Study deep maxout networks and their equivalence to Gaussian processes.
problem Understanding neural networks with infinite width.
method Derive equivalence between deep maxout networks and Gaussian processes, characterize maxout kernel, and provide efficient numerical implementation.
result Bayesian inference based on deep maxout network kernel leads to competitive results compared to finite-width counterparts and deep neural network kernels.
Study examines dependence properties of Bayesian neural network units in finite-width networks.
problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.
Paper characterizes gradient descent dynamics for neural networks with finite width.
problem Characterize gradient descent dynamics for multi-layer neural networks.
method Non-asymptotic state evolution theory for finite-width networks.
result Gradient descent dynamics provide precise distributional characterization.
Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.
problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.
Study infinite-depth limits of neural networks with fixed width.
problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.
Bayesian neural networks approximate Student-t processes in the infinite-width limit.
problem Modeling uncertainty in neural networks with greater flexibility.
method Extending asymptotic properties of Gaussian processes to Student-t processes in the infinite-width limit of BNNs.
result Posterior BNNs converge to Student-t processes in the infinite-width limit.
Study shows bound on Uryson width for specific 3D manifolds.
problem Bounding Uryson width for 3D manifolds with non-negative Ricci curvature and strictly mean convex boundary.
method Proved existence of a Morse function with uniform diameter bounds on level sets.
result Upper bound on Uryson width for the specified 3D manifolds.
This paper studies large-width asymptotics for ReLU neural networks with α-Stable initializations.
problem Characterizing the large-width behavior of ReLU neural networks with α-Stable initializations.
method Analysis of the large-width distributions and training dynamics of ReLU neural networks initialized with α-Stable distributions.
result For ReLU neural networks with α-Stable initializations, the large-width training dynamics achieve zero training error at a linear rate, characterized by a random kernel.
Study Gaussian approximation for deep neural networks with random weights.
problem Understanding the distribution of deep neural networks with random weights.
method Established Gaussian approximation bounds in Wasserstein-1 norm.
result Convergence rates of order n−(1/6)L−1+ε for deep networks with proportional layer widths. In order to investigate the origin of large price fluctuations, we analyze stock price changes of ten frequently traded NASDAQ stocks in the year 2002. Though the influence of the trading frequency on the aggregate return in a certain time interval is important, it cannot alone explain the heavy tailed distribution of …
Uniform convergence of interpolators proven for Gaussian data.
problem Interpolation learning in high-dimensional linear regression with Gaussian data.
method Generic uniform convergence guarantee in terms of Gaussian width.
result Consistency of interpolators for minimum-norm and near-minimal-norm cases.
Study on predicting sequences with Gaussian constraints, linking to intrinsic volumes and metric complexity.
problem Predicting sequences almost as well as the best Gaussian distribution with mean in a given subset.
method Expressed minimax regret in terms of intrinsic volumes, established comparison inequality for Wills functional, characterized global covering numbers and local Gaussian widths.
result Sharp estimates on the log-Laplace transform of intrinsic volume sequence for a general nonconvex set.
The paper extends infinite-width analysis to neural network Jacobians, revealing convergence to Gaussian processes and linear ODEs.
problem Understanding the training dynamics of neural networks in the infinite-width limit.
method Extending infinite-width analysis to Jacobians, characterizing convergence to Gaussian processes and linear ODEs.
result The evolution of MLPs under robust training in the infinite-width limit is described by a linear ODE.
Study of deep neural networks with dependent weights leading to new model limits and properties.
problem Characterizing deep neural networks with dependent weights in the infinite-width limit.
method Modeling weights as a mixture of Gaussian distributions and analyzing the infinite-width limit.
result Characterization of neural network layers by scalar parameters and Lévy measures, leading to new model limits.
Finite-width neural networks are approximated by Gaussian processes with finite size corrections.
problem Understanding the behavior of finite-width neural networks as they approach infinite width.
method Analyzing the distribution of outputs at initialization for large, finite neural networks with a single hidden layer.
result The distribution of outputs at initialization is well described by a Gaussian perturbed by the fourth Hermite polynomial, with the perturbation scale inversely proportional to the number of network units.
Wider neural networks perform better than deeper ones with the same number of parameters.
problem Understanding the role of network width versus the number of parameters in neural network performance.
method Comparing models with different ways of increasing width while keeping the number of parameters constant, analyzing their performance and using Gaussian Process kernels for analysis.
result Network width is the determining factor for good performance, while the number of weights is secondary as long as trainability is ensured.
The paper studies deep neural networks with Gaussian weights and finds their asymptotic behavior.
problem Understanding the behavior of deep neural networks with large width.
method Function-space perspective, Gaussian process analysis, weak convergence in large-width limit.
result Deep neural networks with large width converge to a continuous Gaussian process.
Random neural networks with ReLU activations are non-Gaussian processes.
problem Understanding the behavior of neural networks with random initialization and rectified linear units.
method Proving these networks are non-Gaussian processes and deriving their properties.
result These networks can converge to non-Gaussian processes under certain conditions.
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
New ReLU initialization improves network performance and dynamical isometry.
problem Improving the initialization of ReLU units for better network performance.
method Derive exact joint signal output distribution for fully-connected networks with Gaussian weights and biases, and propose a new initialization scheme for ReLU units.
result Proposed initialization scheme achieves dynamical isometry, improving network performance.
The paper provides non-asymptotic Edgeworth expansions for neural network outputs.
problem Approximating deviations of finite-width neural networks from their Gaussian limit.
method Multidimensional Edgeworth expansions of arbitrary order for neural network outputs.
result Established a bound on the total variation distance between neural network output and its Edgeworth approximation.
Study shows polynomial-width neural networks can closely approximate infinite-width networks in polynomial time.
problem Approximating dynamics of polynomial-width neural networks with infinite-width networks.
method Bounding approximation gap through a differential equation governed by mean-field dynamics, considering local Hessian.
result Polynomially many neurons are sufficient to closely approximate mean-field dynamics.
This paper quantifies how well random neural networks can approximate continuous functions.
problem Approximating continuous functions with random neural networks.
method Investigates three types of random neural networks: infinite width, subsampled, and corrected. Analyzes approximation rates and provides bounds.
result A function can be approximated with complexity proportional to δ and d. Neural network-based post-processing improves ensemble forecast sharpness
problem Reducing the width of central prediction intervals in ensemble forecasts
method Extending loss function with a penalty term
result 8.2%-12.5% reduction in width of central prediction interval
New framework for understanding infinite-width neural networks.
problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.
Analyzes SGD dynamics in two-layer networks, bridging different regimes.
problem Understanding SGD dynamics in high-dimensional and mean-field settings.
method Rigorous analysis via deterministic low-dimensional description of sufficient statistics.
result Infinite-width dynamics remains close to a low-dimensional subspace.
Study shows how large neural networks avoid overfitting through decoupling of feature learning and complexity growth.
problem Understanding inductive bias and generalization in large neural networks.
method Dynamical mean field theory applied to large two-layer networks.
result Training dynamics of large networks exhibit a separation of timescales, decoupling feature learning and overfitting.
Sharp risk bounds for early-stopping in Gaussian linear regression are derived.
problem Minimizing in-sample mean squared error in high-dimensional Gaussian linear regression.
method Early-stopped mirror descent (ESMD) with local Gaussian width bounds.
result Sharp risk bounds extend to early-stopped mirror descent for least squares estimator (LSE).
The oloid is the convex hull of two circles with equal radius in perpendicular planes so that the center of each circle lies on the other circle. We calculate the mean width of the oloid in two ways, first via the integral of mean curvature, and then directly. Using this result, the surface area and the volume of the p…