The paper studies ReLU layers, introducing new tools to understand their singular values and Gaussian mean width.
problem Understanding the role of ReLU layers in neural networks and their impact on network performance.
method Introducing ReLU singular values and Gaussian mean width of operators to study ReLU layers.
result ReLU singular values and Gaussian mean width provide metrics for distinguishing correctly and incorrectly classified data.
PHP connects to ReLU neural networks for scalable Bayesian inference.
problem Scalability and Bayesian inference in two-layer ReLU neural networks.
method PHP with Gaussian prior, decomposition propositions, annealed sequential Monte Carlo.
result PHP provides an alternative scalable representation for two-layer ReLU neural networks.
Gradient descent converges to minimum Bayes risk for two-layer ReLU networks in mean field regime.
problem Training two-layer ReLU networks using gradient descent in the mean field regime.
method Describes a condition for convergence to minimum Bayes risk, extending previous results to ReLU-activated networks.
result The condition for convergence does not depend on initialization and concerns weak convergence of network realization.
Analytical method finds deeper optima in two-layer ReLU networks.
problem Training two-layer ReLU networks with analytical methods.
method Analytically finding critical points of the loss function for one layer while keeping the other fixed.
result Significantly smaller training loss values on real datasets compared to gradient descent methods.
Polynomial-time convex optimization for CNNs with ReLU activations.
problem Training Convolutional Neural Networks (CNNs) with ReLU activations.
method Developed a convex analytic framework using semi-infinite duality to formulate equivalent convex optimization problems for CNN architectures.
result Proved that two-layer CNNs can be globally optimized via an ℓ 2 \ell_2 ℓ 2 norm regularized convex program. A new activation function FTS improves deep learning performance.
problem Hindered propagation of negative values in ReLU.
method Proposed Flatten-T Swish (FTS) activation function, evaluated on MNIST dataset.
result FTS with T=-0.20 improves MNIST classification accuracy by 1.15% on 8-layer DFNN.
ReLU units can 'die' in neural networks, causing slower convergence.
problem ReLU units sometimes produce near-zero outputs during training.
method Simulation and statistical analysis of a simplified ReLU unit model.
result Activation probability decreases as training progresses, leading to slower convergence.
Deep ReLU networks can be simplified to a three-layer model.
problem Understanding the behavior of deep neural networks.
method Constructive proof and algorithm to transform deep networks into shallow ones.
result Deep ReLU networks can be represented by a simpler three-layer structure.
Gradient descent learns ReLU networks with Gaussian inputs and noisy outputs.
problem Learning one-hidden-layer ReLU networks with Gaussian inputs and noisy outputs.
method Gradient descent with tensor initialization for empirical risk minimization.
result Gradient descent converges to ground-truth parameters at a linear rate up to statistical error.
The paper proves neural networks with ReLU and softmax can approximate any function.
problem Approximating functions and class labels in neural networks.
method Extended universal approximator theory to neural networks with ReLU and softmax.
result Neural networks with ReLU and softmax can approximate any function and class labels.
Estimates generalization error for two-layer ReLU NNs through minimum norm solutions.
problem Estimating generalization error for two-layer ReLU NNs trained by mean squared error.
method Uses minimum norm solutions and Neural Tangent Kernel (NTK) regime to derive generalization error bounds.
result Derives an a priori generalization error bound for two-layer ReLU NNs without requiring exponentially large number of neurons.
The paper calculates bounds on the local Lipschitz constants of neural network layers.
problem Understanding the Lipschitz constants of neural network layers for robustness analysis.
method Analytical approach to determine upper bounds on local Lipschitz constants of affine-ReLU functions.
result The method produces tighter bounds than the standard conservative bound, especially for small perturbations.
Random ReLU features are shown to be a universally consistent learning algorithm but struggle with complex functions.
problem Approximating complex functions with random ReLU features.
method Study of random ReLU features through their RKHS and composition of functions.
result Random ReLU features can efficiently approximate complex functions but not as well as multi-layer ReLU networks.
Deep ReLU networks show that 4 layers suffice for unique input recovery.
problem Injectivity capacity of deep ReLU networks.
method Developed a program connecting deep ReLU injectivity to an l l l -extension of the ℓ 0 \ell_0 ℓ 0 spherical perceptrons, using random duality theory. result Only 4 layers are needed for unique input recovery, showing expansion saturation effect.
Researchers create integral representations for two-layer ReLU networks with quantitative bounds.
problem Approximating functions with two-layer ReLU networks using explicit integral representations.
method Developed integral representations involving harmonic extension and projection, providing L 2 L^{2} L 2 bounds. result Functions can be approximated with L 2 L^{2} L 2 errors independent of dimension or degree, depending on coefficients and distribution. This paper analyzes how normalization layers improve neural network training.
problem Improving generalization performance and training speed of neural networks.
method Global convergence analysis of two-layer neural networks with ReLU activations and Weight Normalization.
result Introduction of normalization layers changes the optimization landscape, enabling faster convergence.
ReLU networks trained with MILPs match deep learning accuracy.
problem Training deep neural networks efficiently.
method Iterative training with Mixed Integer Linear Programs (MILPs).
result ReLU networks can be trained with MILPs achieving similar accuracy to deep learning methods.
Proves SQ lower bounds for learning two-hidden-layer neural networks.
problem Learning two-hidden-layer ReLU networks with Gaussian inputs.
method Refined lifting procedure to reduce Boolean PAC learning to Gaussian.
result Superpolynomial SQ lower bounds for Gaussian inputs.
The study examines Fisher information matrices and neural tangent kernels for simple ReLU networks with random weights.
problem Understanding the relationship between Fisher information matrices and neural tangent kernels for 2-layer ReLU networks.
method Analyzes Fisher information matrices and neural tangent kernels for 2-layer ReLU networks with random hidden weights, focusing on spectral decomposition and eigenfunctions.
result Obtained an approximation formula for functions represented by 2-layer neural networks.
Two-layer ReLU networks can overfit without harm, study finds.
problem Understanding when and how two-layer ReLU networks can overfit without harming generalization.
method Established algorithm-dependent risk bounds for two-layer ReLU convolutional neural networks with label-flipping noise.
result Gradient descent-trained ReLU networks can achieve near-zero training loss and Bayes optimal test risk.
New Banach spaces for ReLU networks enable better function approximation and gradient dynamics analysis.
problem Function approximation and gradient dynamics in multi-layer ReLU networks.
method Developed Banach spaces for ReLU networks, defined new function representations, and analyzed gradient flow dynamics.
result Gradient flow dynamics of the new representation is the continuous analog of gradient descent for ReLU networks.
The study examines if ReLU activation function is optimal for modularity in neural networks.
problem Finding the best activation function for modularity in neural networks.
method Comparing ReLU with other activation functions for modularity and performance.
result ReLU may not be the best choice for modularity, suggesting other functions could be more suitable.
Study reveals properties of local minima in ReLU networks.
problem Understanding the loss landscape of neural networks.
method Theoretical analysis of one-hidden-layer ReLU networks.
result All differentiable local minima are global within certain regions.
New approach replaces BN layers with shifted-ReLU for embedded systems.
problem Complexity and overheads of BN layers in low-power embedded systems.
method Used shifted-ReLU layers instead of BN layers in wide residual networks.
result Shifted-ReLU layers offer advantages in speed, memory, and complexity without significant accuracy loss.
The study characterizes sets for which one-layer neural networks are positive.
problem Characterizing sets of points for which one-layer neural networks are positive.
method Investigation of subsets of the real plane for one-layer ReLU neural networks.
result Full characterization of such sets for cones, and necessary condition for any subset of \(\mathbb{R}^d\).
The paper shows how random ReLU networks converge to smooth splines.
problem Understanding the behavior of shallow ReLU neural networks with random weights.
method Mathematical analysis of L2-regularized regression and gradient descent.
result Random ReLU networks converge to smooth splines as the number of hidden nodes increases.
Adversarial examples in deep ReLU networks with constant depth.
problem Adversarial perturbations lead to large output changes in deep ReLU networks.
method Analyzes the phenomenon in networks with independent Gaussian parameters and constant depth.
result Adversarial examples arise due to functions being close to linear.
Gradient descent proves global convergence for deep networks with a single wide layer.
problem Proving global convergence of gradient descent for deep ReLU networks.
method Simplified proof using a single wide layer, leveraging ReLU's Lipschitz property.
result Gradient descent converges globally for networks with a single wide layer.
Two-layer CNNs can overfit well if initialized correctly.
problem Understanding the conditions for benign overfitting in over-parameterized CNNs.
method Extending analysis to fully trainable two-layer CNNs, examining initialization scaling effects.
result Initialization scaling of the output layer is crucial; large scales lead to fixed output behavior, small scales to complex interactions.
The study calculates the injectivity capacity of ReLU networks using a novel mathematical approach.
problem Determining the injectivity capacity of ReLU networks layers.
method Employing fully lifted random duality theory (fl RDT) to handle the ℓ 0 \ell_0 ℓ 0 spherical perceptron and implicitly the ReLU layers injectivity. result The lifting mechanism converges remarkably fast with relative corrections not exceeding 0.1%.
In this paper we investigate the family of functions representable by deep neural networks (DNN) with rectified linear units (ReLU). We give an algorithm to train a ReLU DNN with one hidden layer to *global optimality* with runtime polynomial in the data size albeit exponential in the input dimension. Further, we impro…
Student network specializes teacher nodes in deep ReLU networks.
problem Training deep ReLU networks with finite width and input dimension.
method Stochastic Gradient Descent (SGD) on over-realized student network trained from teacher network output.
result Each teacher node is specialized by at least one student node at the lowest layer under mild conditions.
Least symmetry breaking principle explains SGD's local minima in shallow ReLU networks.
problem Understanding the structure of local minima in two-layer ReLU networks.
method Analyzing the squared loss optimization problem for ReLU networks with Gaussian inputs and applying the principle of least symmetry breaking.
result The principle of least symmetry breaking explains the structure of spurious local minima detected by SGD.
For any positive integer k k k , there exist neural networks with Θ ( k 3 ) Θ(k^3) Θ ( k 3 ) layers, Θ ( 1 ) Θ(1) Θ ( 1 ) nodes per layer, and Θ ( 1 ) Θ(1) Θ ( 1 ) distinct parameters which can not be approximated by networks with O ( k ) \mathcal{O}(k) O ( k ) layers unless they are exponentially large --- they must possess Ω ( 2 k ) Ω(2^k) Ω ( 2 k ) nodes. This result is proved here for a class o…
3-layer ReLU networks with Ω(√N) nodes can memorize most datasets.
problem Understanding memorization capacity of ReLU networks.
method Analyzing depth and width requirements for memorization.
result Width Θ(√N) is necessary and sufficient for memorizing N data points.
New algorithm for learning ReLU networks with Gaussian noise, improving previous results.
problem PAC learning one-hidden-layer ReLU networks with Gaussian marginals and label noise.
method First polynomial-time algorithm for k k k up to i l d e O ( log d ) ilde{O}(\sqrt{\log d}) i l d e O ( log d ) for positive coefficients, no assumptions on rank or condition number. result Proves a Statistical Query lower bound of d Ω ( k ) d^{Ω(k)} d Ω ( k ) for arbitrary real coefficients, separating learnability classes. Stable unactivated neurons reduce expressiveness in ReLU networks.
problem Reducing expressiveness in ReLU neural networks due to stably unactivated neurons.
method Investigated the probability of neurons being stably unactivated in ReLU networks with symmetric weight and bias distributions.
result Proved the probability of a neuron being stably unactivated in the second hidden layer of a ReLU network.
Study shows trained neural networks can overfit without bias or variance issues.
problem Understanding overfitting in trained two-layer ReLU networks.
method Analysis of gradient flow in the neural tangent kernel regime, decomposition of excess risk.
result Trained networks can overfit benignly without bias or variance issues.
New properties of deep ReLU networks at initialization improve performance.
problem Improving performance of deep ReLU networks.
method PAC analysis framework to prove novel properties of He initialization.
result Hidden activation norms and weight gradient norms are preserved under He initialization.
New algorithm trains ReLU networks via alternating minimization.
problem Training deep neural networks with ReLU activations.
method Alternating minimization of activation patterns and weight updates.
result Proves linear convergence for recovering true parameters.
This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed d i n ≥ 1 , d_{in}\geq 1, d in ≥ 1 , what is the minimal width w w w so that neural nets with ReLU activations, input dimension d i n d_{in} d in , hidden layer widths at most w , w, w , and …
Gradient descent fails to train two-layer ReLU networks, leading to poor performance.
problem Gradient descent training of two-layer ReLU networks initialized by He et al. (2015) fails to find optimal solutions.
method Gradient descent on a least-squares loss for training two-layer (Leaky)ReLU networks.
result Gradient descent only finds bad local minima, leading to linear regression for non-linear target functions.
The paper analyzes FGSM and CW-L2 attacks on CNNs.
problem Efficacy of adversarial attacks on neural networks.
method Theoretical analysis and numerical verification of FGSM and CW-L2 attacks.
result Theoretical findings on the effectiveness of FGSM and CW-L2 attacks on CNNs.
A new Shapley value approach for neural networks interpretable and stable.
problem Neural networks' interpretability and training stability issues.
method Shapley value approximation for ReLU activation, globally continuous Shapley gradient, Shapley Activation function.
result SA consistently outperforms ReLU in training convergence, accuracy, and stability.
Paper finds conditions for benign overfitting in neural networks.
problem Benign overfitting in leaky ReLU two-layer neural networks.
method Established directional convergence and classification error bounds.
result Benign overfitting occurs with high probability on mixture data.
We solve the optimization of two-layer ReLU networks using convex math.
problem Optimizing two-layer ReLU neural networks.
method Exact characterization of optimal solutions via convex optimization.
result We prove that all globally optimal solutions can be found via convex optimization.
Study of two-layer ReLU neural network phase diagram at infinite-width limit.
problem Characterize the dynamical regimes of two-layer ReLU neural networks.
method Combining experimental and theoretical approaches, including phase diagram analogy.
result Identification of three regimes: linear, critical, and condensed.
Study analyzes error in ReLU networks with local connections.
problem Improving neural network performance and understanding approximation errors.
method Analyzed approximation error of ReLU networks with local connections.
result Error estimate depends on depth and width of hidden layers.