Critical initialisation strategies are identified for noisy ReLU networks.
problem Understanding signal propagation in noisy rectifier neural networks.
method Developed a new framework for signal propagation in stochastic regularized neural networks, incorporating various noise distributions.
result Critical initialisation strategies for multiplicative noise (e.g. dropout) are identified, but not for additive noise.
New hypothesis and method reduce noise in saliency maps.
problem Noisy saliency maps from deep neural networks.
method Proposed a new hypothesis and Rectified Gradient method.
result Rectified Gradient effectively reduces saliency map noise.
Neural networks can approximate rectifiable measures with small error.
problem Approximating complex rectifiable measures using neural networks.
method Using ReLU neural networks to approximate (countably) m-rectifiable measures as push-forwards of the Lebesgue measure.
result The approximation error in terms of Wasserstein distance can be made arbitrarily small.
In this paper, we introduce transformations of deep rectifier networks, enabling the conversion of deep rectifier networks into shallow rectifier networks. We subsequently prove that any rectifier net of any depth can be represented by a maximum of a number of functions that can be realized by a shallow network with a …
We study the problem of learning one-hidden-layer neural networks with Rectified Linear Unit (ReLU) activation function, where the inputs are sampled from standard Gaussian distribution and the outputs are generated from a noisy teacher network. We analyze the performance of gradient descent for training such kind of n…
In this paper we investigate the performance of different types of rectified activation functions in convolutional neural network: standard rectified linear unit (ReLU), leaky rectified linear unit (Leaky ReLU), parametric rectified linear unit (PReLU) and a new randomized leaky rectified linear units (RReLU). We evalu…
Paper proposes a new activation function to reduce overfitting and large weight update issues.
problem Overfitting and large weight update problems in neural networks.
method Introduces a new activation function called Thresholded Exponential Rectified Linear Units (TERELU).
result TERELU shows better performance in reducing overfitting and large weight update issues compared to other activation functions.
Random neural networks produce functions with a number of knots equal to the number of neurons.
problem Understanding why neural networks with many parameters do not overfit early in training.
method Analyzing random scalar-input feed-forward rectified linear unit architectures, showing they are random linear splines.
result The number of knots in random neural networks is equal to the number of neurons, to very close approximation.
CRITS improves time series classification with interpretable local explanations.
problem Lack of detailed explanations in time series classification models.
method CRITS uses convolutional kernels, max-pooling, and rectified linear units to extract feature weights.
result CRITS provides intrinsically interpretable local explanations without requiring gradients or random perturbations.
This paper provides a theoretical justification of the superior classification performance of deep rectifier networks over shallow rectifier networks from the geometrical perspective of piecewise linear (PWL) classifier boundaries. We show that, for a given threshold on the approximation error, the required number of b…
Methods from convex optimization are widely used as building blocks for deep learning algorithms. However, the reasons for their empirical success are unclear, since modern convolutional networks (convnets), incorporating rectifier units and max-pooling, are neither smooth nor convex. Standard guarantees therefore do n…
Dropout training improves neural networks' performance.
problem Improving neural network convergence and generalization.
method Two-layer neural networks with ReLU activations, overparametrization, and positive margin assumption.
result Dropout training achieves ε-suboptimality in test error in O(1/ε) iterations.
Rectified flows achieve optimal sample complexity for generating data.
problem Generating high-quality data samples efficiently.
method Rectified flows constrain transport trajectories to be linear, enabling efficient sampling.
result Achieve sample complexity of ildeO(ε−2), matching optimal rate for mean estimation. This paper controls the capacity of weight-normalized deep neural networks using rectified linear units.
problem Capacity control of weight-normalized deep neural networks.
method Establishes upper bounds on Rademacher complexities and analyzes approximation properties of Lp,q weight normalized networks. result For L1,∞ weight normalized networks, the approximation error is controlled by the L1 norm of the output layer, and generalization error depends on the square root of depth. Improved texture synthesis using wavelet-based statistics with rectifier non-linearity.
problem Improving texture synthesis quality using wavelet representations.
method Proposes a family of statistics based on non-linear wavelet representations with a generalized rectifier non-linearity.
result Significantly improves visual quality of texture synthesis compared to classical wavelet-based models.
Improved activation function NLReLU boosts neural network performance.
problem Performance issues with ReLU activation function.
method NLReLU uses parametric natural logarithmic transform to improve ReLU.
result NLReLU provides higher accuracy than ReLU in various neural networks.
We consider a neural network architecture with randomized features, a sign-splitter, followed by rectified linear units (ReLU). We prove that our architecture exhibits robustness to the input perturbation: the output feature of the neural network exhibits a Lipschitz continuity in terms of the input perturbation. We fu…
AReLU uses attention-based rectification to improve neural network performance.
problem Improving neural network performance through better activation functions.
method Integrates attention mechanism with rectified linear unit (ReLU) to learn and scale feature maps.
result AReLU significantly boosts performance of most network architectures with minimal changes.
Empirical bounds estimate the number of linear regions in deep ReLU networks.
problem Estimating the number of linear regions in deep neural networks.
method Empirical bounds based on activation patterns and probabilistic inference.
result Fast proxy for the number of linear regions of deep neural networks.
Modern convolutional networks, incorporating rectifiers and max-pooling, are neither smooth nor convex; standard guarantees therefore do not apply. Nevertheless, methods from convex optimization such as gradient descent and Adam are widely used as building blocks for deep learning algorithms. This paper provides the fi…
New framework explains deep neural networks using variational spline theory.
problem Understanding functions learned by deep neural networks.
method Developed a variational framework and function space.
result Deep ReLU networks are solutions to regularized data fitting problems over the proposed function space.
A neural network with a single hidden layer can't represent certain multivariable functions.
problem Representing certain multivariable functions with a neural network having only one hidden layer.
method Developed a continuum version of a one-hidden-layer neural network with ReLU activation, and proved constraints on its parameters and second derivative.
result Existence of a smooth binary function that cannot be precisely represented by any such neural network.
New method trains deep vanilla networks as fast as ResNets without shortcut connections.
problem Training very deep neural networks is challenging.
method Developed a new type of transformation compatible with Leaky ReLUs.
result Validation accuracies with deep vanilla networks are competitive with ResNets and significantly higher.
We propose to model the acoustic space of deep neural network (DNN) class-conditional posterior probabilities as a union of low-dimensional subspaces. To that end, the training posteriors are used for dictionary learning and sparse coding. Sparse representation of the test posteriors using this dictionary enables proje…
We introduce a new neural network model, together with a tractable and monotone online learning algorithm. Our model describes feed-forward networks for classification, with one output node for each class. The only nonlinear operation is rectification using a ReLU function with a bias. However, there is a rectifier on …
Scattering networks maximize separation on low-dimensional data.
problem Maximizing separation capacity on low-dimensional datasets.
method Characterize and bound separation capacity for feature extractors, then apply to scattering networks with specific criteria.
result Design criteria for scattering networks to maximize separation on low-dimensional data.
Study rectifying curves in 3D multiplicative Euclidean space.
problem Investigate rectifying curves in a non-Newtonian geometry setting.
method Apply multiplicative differential-geometric concepts to rectifying curves.
result Classify multiplicative rectifying curves using spherical curves.
Investigates Darboux rectifying curves on smooth surfaces.
problem Characterizing Darboux rectifying curves on smooth surfaces.
method Analyzes the position vector under isometry and finds conformal invariance conditions.
result Identifies sufficient conditions for conformal invariance of Darboux rectifying curves.
A space curve in a Euclidean 3-space E3 is called a rectifying curve if its position vector field always lies in its rectifying plane. This notion of rectifying curves was introduced by the author in [Amer. Math. Monthly {\bf 110} (2003), no. 2, 147-152]. In this present article, we introduce and study the n…
Study characterizes k-rectifiable sets in homogeneous groups.
problem Characterizing k-rectifiable sets in arbitrary homogeneous groups. method Proves characterizations using (k,G)-approximate tangent groups. result Existence of (k,G)-approximate tangent groups implies k-rectifiability. New curves generalize helix and rectifying curves.
problem Generalizing helix and rectifying curves.
method Introducing f-rectifying curves with f-position vector in rectifying plane.
result Classification and characterization of f-rectifying curves.
This study uses neural networks to solve interpolation problems with sparse, infinitely wide layers.
problem Exact data interpolation using sparse, infinitely wide neural networks.
method Atomic norm framework to derive convex hulls and equivalent convex formulations.
result Simple characterizations of convex hulls for different constraints on network weights and biases.
Recalls and refines the concept of algebraically rectifiable curves.
problem Classical notion of algebraically rectifiable plane curves.
method Provides new criteria, relates to quadratic differentials, and generalizes to higher order differentials.
result Generalization and new criteria for algebraic rectifiability.
A model of associative memory is studied, which stores and reliably retrieves many more patterns than the number of neurons in the network. We propose a simple duality between this dense associative memory and neural networks commonly used in deep learning. On the associative memory side of this duality, a family of mo…
Harmonic maps to Euclidean buildings have rectifiable singular strata.
problem Understanding the structure of singular points for harmonic maps.
method Defining singular strata and proving rectifiability using the rectifiable Reifenberg program.
result Rectifiability of singular strata for harmonic maps into F-connected complexes. We present QuickNet, a fast and accurate network architecture that is both faster and significantly more accurate than other fast deep architectures like SqueezeNet. Furthermore, it uses less parameters than previous networks, making it more memory efficient. We do this by making two major modifications to the referenc…
Uniform rectifiability proven for sets with Poincaré inequalities.
problem Uniform rectifiability of sets with Poincaré inequalities.
method Weak (1,d)-Poincaré inequality and surface measure. result Uniform rectifiability achieved for sets supporting such inequalities.
We prove (without using Federer's structure theorem) that a finite-mass flat chain over any coefficient group is rectifiable if and only if almost all of its 0-dimensional slices are rectifiable. This implies that every flat chain of finite mass and finite size is rectifiable. It also leads to a simple necessary and su…
The paper analyzes deep neural networks using rectified linear units.
problem Understanding the individual affine linear representations of deep neural networks.
method Signal processing perspective, atomic decompositions, Lipschitz regularity estimation.
result Conditions for stabilizing learning in deep neural networks without network depth constraints.
Estimates parameters of a rectified Gaussian distribution using ReLU networks.
problem Estimating parameters of a rectified Gaussian distribution from i.i.d. samples.
method Simple algorithm using O(1/ε2) samples and O(d2/ε2) time. result Estimates distribution up to ε in total variation distance. New methods for solving hydrodynamic-type equations using quasi-rectifiable Lie algebras.
problem Solving systems of hydrodynamic-type equations.
method Introducing and analyzing quasi-rectifiable Lie algebras and vector fields.
result New methods for solving hydrodynamic-type equations.
Rectifies singular set of harmonic maps into complex.
problem Regularity of harmonic maps into complex manifolds.
method Proves (m−2)-rectifiability of singular set. result Singular set is (m−2)-rectifiable. The paper generalizes rectifying and normal curves in Lorentzian n-space.
problem Characterizing and classifying g−rectifying and g−normal curves in Lorentzian n-space. method Introducing a g−position vector field and defining g−rectifying and g−normal curves based on this field. result Comprehensive characterization and classification of g−rectifying and g−normal curves. Study approximates nonlinear functionals using deep ReLU networks.
problem Approximating nonlinear continuous functionals with neural networks.
method Constructs continuous piecewise linear interpolation under simple triangulation, analyzes rates of approximation.
result Established rates of approximation for functional deep ReLU networks.
We show that C^1 hypersurfaces in the Heisenberg group are countably N-rectifiable. As a corollary, this shows that all C^1_H graphs over the xy-plane are countable N-rectifiable, showing the equivalence of this notion of rectifiability with that of Franchi, Serra Cassano and Serapioni for such surfaces.
The paper characterizes special curves and generalizes rectifying-type curves in n-dimensional space.
problem Characterizing and generalizing special curves in higher dimensions.
method Characterization through Rotation minimizing frame (RMF) and generalization of rectifying-type curves.
result Rectifying-type curves are generalized in n-dimensional space.
New algorithm reveals piecewise affine structure of neural networks.
problem Lack of strong guarantees on deep neural networks' behavior in safety-critical applications.
method Developed a novel algorithm to compute the piecewise affine form of neural networks.
result Computed piecewise affine representations of neural networks with rectified linear unit activations.
The notion of rectifying curve in the Euclidean space is introduced by Chen as a curve whose position vector always lies in its rectifying plane spanned by the tangent and the binormal vector field t and n_2 of the curve. In this study, we have obtained some characterizations of semi-real spatial quaternionic rectifyin…