New insights into neural network training efficiency.
problem Understanding the optimal initialization for deep neural networks.
method Exploring the edge of chaos and saturation of tanh activation function.
result The line of uniformity in phase space intersects the edge of chaos, indicating saturation begins to hinder training efficiency.
It has long been suggested that the biological brain operates at some critical point between two different phases, possibly order and chaos. Despite many indirect empirical evidence from the brain and analytical indication on simple neural networks, the foundation of this hypothesis on generic non-linear systems remain…
The paper analyzes deep neural networks' expressivity and training, revealing critical expressivity issues.
problem Critical expressivity issues in deep neural networks.
method Quantitative analysis using Hilbert space and Hermite polynomials for feature mapping and activation function design.
result Deep neural networks evolve to the edge of chaos, but expressivity depends on overcoming convergence.
Herding defines a deterministic dynamical system at the edge of chaos. It generates a sequence of model states and parameters by alternating parameter perturbations with state maximizations, where the sequence of states can be interpreted as "samples" from an associated MRF model. Herding differs from maximum likelihoo…
Optimizes wide low-rank neural networks for reduced parameters and cost.
problem Reducing the number of learnable parameters in wide neural networks.
method Analyzed edge-of-chaos dynamics and derived formulae for optimal weight and bias variances.
result Optimal weight and bias variances for low-rank networks follow from multiplicative scaling.
The weight initialization and the activation function of deep neural networks have a crucial impact on the performance of the training procedure. An inappropriate selection can lead to the loss of information of the input during forward propagation and the exponential vanishing/exploding of gradients during back-propag…
Dropout schedules can be optimized to significantly reduce model test loss.
problem Improving model performance in neural networks.
method Developed a mean-field theory of dropout at the edge of chaos, proposing front-loaded dropout schedules.
result Front-loaded dropout schedules reduce test loss by 18-35% over constant dropout.
Study challenges the Gaussian pre-activations assumption in neural networks.
problem Challenges the assumption that pre-activations are Gaussian in neural networks.
method Constructs pairs of activation functions and initialization distributions to ensure Gaussian pre-activations.
result Discovered constraints for ensuring Gaussian pre-activations in neural networks.
The weight initialization and the activation function of deep neural networks have a crucial impact on the performance of the training procedure. An inappropriate selection can lead to the loss of information of the input during forward propagation and the exponential vanishing/exploding of gradients during back-propag…
This work bridges theory and practice in spiking reservoirs, identifying robust parameter ranges.
problem Challenging tuning of spiking reservoirs at the edge-of-chaos.
method Introducing robustness interval, systematic evaluations, and control experiments.
result Consistent monotonic trends in robustness interval width across network configurations.
Deep neural networks near edge of chaos show universal scaling laws.
problem Understanding the behavior of deep neural networks near critical points.
method Analogy to absorbing phase transitions in statistical mechanics, deterministic propagation dynamics, mean-field and directed percolation universality classes.
result Deep neural networks exhibit universal scaling laws near the edge of chaos.
We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output Jacobian of N is exponential in a simple architecture-dependent constant beta, gi…
Develops a new theory for neural systems stability and width effects.
problem Stability and finite-width effects in deep neural systems.
method Gauge-covariant stochastic effective field theory using classical commuting fields.
result Predicts the edge of chaos and low-frequency spectral deformation.
We study the concentration of NTK for MLPs at EOC, proving finite-width approximation of gradient independence.
problem Understanding the concentration of Neural Tangent Kernel (NTK) for MLPs at the Edge of Chaos (EOC).
method Proved approximate gradient independence holds at finite width, using maximal inequalities to show NTK matrix concentrates around its infinitely wide limit.
result The NTK matrix of MLPs at EOC concentrates around its infinitely wide limit, requiring hidden layer widths to grow quadratically.
We study the behavior of untrained neural networks whose weights and biases are randomly distributed using mean field theory. We show the existence of depth scales that naturally limit the maximum depth of signal propagation through these random networks. Our main practical result is to show that random networks may be…
New method predicts neural network performance using free probability theory.
problem Stability and performance prediction of feed-forward neural networks.
method Free Probability Theory and homotopy method for Jacobian spectral density computation.
result FPT metrics correlate highly with final test accuracies of neural networks.
Study on neural networks with non-normal interactions reveals unique spectral properties.
problem Understanding episodic memory encoding in the brain.
method Developed a neural network model with non-Hermitian couplings and applied random matrix theory.
result Spectral density of the model is non-uniform and can transition to chaos, providing computational benefits.
This paper extends the Gaussian process interpretation of deep networks to more varied weight distributions.
problem Understanding the impact of different weight initialization schemes on deep learning dynamics.
method Extending the Gaussian process interpretation to PSEUDO-IID weight distributions, including sparse and low-rank networks.
result PSEUDO-IID initialized networks are effectively equivalent up to variance, enabling tractable posterior distributions.
New method trains deep vanilla networks as fast as ResNets without shortcut connections.
problem Training very deep neural networks is challenging.
method Developed a new type of transformation compatible with Leaky ReLUs.
result Validation accuracies with deep vanilla networks are competitive with ResNets and significantly higher.
Quantum field theory connects deep neural networks to criticality.
problem Understanding the criticality and training dynamics of deep neural networks.
method Constructing quantum field theory for deep neural networks, computing corrections to correlation functions.
result Found precise analogy with O(N) vector model, providing corrections to correlation length. Despite the widespread practical success of deep learning methods, our theoretical understanding of the dynamics of learning in deep neural networks remains quite sparse. We attempt to bridge the gap between the theory and practice of deep learning by systematically analyzing learning dynamics for the restricted case o…