Orthogonal initialization does not speed up training in ultra-wide neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The selection of initial parameter values for gradient-based optimization of deep neural networks is one of the most impactful hyperparameter choices in deep learning systems, affecting both convergence times and model performance. Yet despite significant empirical and theoretical analysis, relatively little has been p…
Initialization of parameters in deep neural networks has been shown to have a big impact on the performance of the networks (Mishkin & Matas, 2015). The initialization scheme devised by He et al, allowed convolution activations to carry a constrained mean which allowed deep networks to be trained effectively (He et al.…
We develop the idea of using an algebraic-geometry approach to classical differential geometry problems. Consider an orthogonal net constructed according to algebraic-geometric data we obtain a set of smooth orthogonal nets that are Ribaucour transformations of the initial orthogonal net.
Optimal spectral estimators and AMP combine for efficient weak recovery in orthogonally invariant GLMs.
Recently mean field theory has been successfully used to analyze properties of wide, random neural networks. It gave rise to a prescriptive theory for initializing feed-forward neural networks with orthogonal weights, which ensures that both the forward propagated activations and the backpropagated gradients are near $…
Deep networks with orthogonal weights show stable fluctuations, improving generalization and training speed.
This paper concerns dictionary learning, i.e., sparse coding, a fundamental representation learning problem. We show that a subgradient descent algorithm, with random initialization, can provably recover orthogonal dictionaries on a natural nonsmooth, nonconvex minimization formulation of the problem, under mi…
New method stabilizes deep neural networks by setting Lyapunov exponent to zero.
A parametric manifold can be viewed as the manifold of orbits of a (regular) foliation of a manifold by means of a family of curves. If the foliation is hypersurface orthogonal, the parametric manifold is equivalent to the 1-parameter family of hypersurfaces orthogonal to the curves, each of which inherits a metric and…
OPT framework improves neural network generalization by learning an orthogonal transformation.
Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients. Short of designing new RNN architectures, previous methods for dealing with this problem usually boil down to orthogonalization of the recurrent dynamics,…
Batch normalization makes deep neural networks' representations increasingly orthogonal.
A machine learning method selects optimal orthonormal bases for functional data analysis.
We solve the equivalence problem for the orthogonally separable webs on the three-sphere under the action of the isometry group. This continues a classical project initiated by Olevsky in which he solved the corresponding canonical forms problem. The solution to the equivalence problem together with the results by Olev…
Study optimizes estimation of orthogonal and rotation matrices from noisy data.
A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We rigorously show almost sure asymptotic freeness of layer-wise Jacobians of deep neura…
Study spectral learning for odeco tensors, addressing initialization bottlenecks.
In our previous article [Rad16], we investigated the asymptotic behaviour of orthogonal Bianchi class B perfect fluids close to the initial singularity and proved the Strong Cosmic Censorship conjecture in this setting. In several of the statements, the case of a stiff fluid had to be excluded. The present paper fills …
The Strong Cosmic Censorship conjecture states that for generic initial data to Einstein's field equations, the maximal globally hyperbolic development is inextendible. We prove this conjecture in the class of orthogonal Bianchi class B perfect fluids and vacuum spacetimes, by showing that unboundedness of certain curv…
New algorithms improve tensor CP decomposition under mild conditions.
Revisits CP tensor decomposition for noisy, non-orthogonal data.
Tensor CANDECOMP/PARAFAC (CP) decomposition is an important tool that solves a wide class of machine learning problems. Existing popular approaches recover components one by one, not necessarily in the order of larger components first. Recently developed simultaneous power method obtains only a high probability recover…
Neural networks learn incrementally from orthogonal data, interpolating with minimal complexity.
It is well known that the initialization of weights in deep neural networks can have a dramatic impact on learning speed. For example, ensuring the mean squared singular value of a network's input-output Jacobian is is essential for avoiding the exponential vanishing or explosion of gradients. The stronger condi…
Study shows RFRR's effectiveness with nearly orthogonal data in overparameterized settings.
In recent years, state-of-the-art methods in computer vision have utilized increasingly deep convolutional neural network architectures (CNNs), with some of the most successful models employing hundreds or even thousands of layers. A variety of pathologies such as vanishing/exploding gradients make training such deep n…
This paper considers the recovery of a rank positive semidefinite matrix from scalar measurements of the form (i.e., quadratic measurements of ). Such problems arise in a variety of applications, including covariance sketching of high-dimensional data…
Adaptive orthogonalization of data for clustering and visualization.
This study explains gradient flow dynamics in neural networks for small initialisation.
Many scientific questions require estimating the effects of continuous treatments. Outcome modeling and weighted regression based on the generalized propensity score are the most commonly used methods to evaluate continuous effects. However, these techniques may be sensitive to model misspecification, extreme weights o…
This paper examines weight initialization for 1-Lipschitz networks to improve robustness against adversarial attacks.
NS-RGS improves orthogonal group synchronization with faster convergence.
Study uses random matrix theory to improve tensor approximation accuracy.
Gradient descent with large steps leads to chaotic parameter space and unpredictable outcomes.
We develop a mean-field theory for multi-component ICA in high dimensions.
Overparameterized models improve performance in sequential learning tasks.
Investigates harmonic self-maps' stability on cohomogeneity one manifolds.
SGD batch size affects autoencoder global minima sparsity and sharpness.
Over a compact oriented manifold, the space of Riemannian metrics and normalised positive volume forms admits a natural pseudo-Riemannian metric , which is useful for the study of Perelman's functional. We show that if the initial speed of a -geodesic is -orthogonal to the tangent space to the or…
Abundant literature has been published on approximation methods for the forward initial margin. The most popular ones being the family of regression methods. This paper describes the mathematical foundations on which these regression approximation methods lie. We introduce mathematical rigor to show that in essence, al…
New method for sampling on constrained domains using orthogonal-space gradient flow.
In this paper we study a model of random knots obtained by fixing a space curve in -dimensional Euclidean space with , and orthogonally projecting the space curve on to random dimensional subspaces. By varying the space curve we obtain different models of random parametrized knots, and we will study how the…
Two-layer networks trained on low-dimensional subspaces are vulnerable to adversarial examples.
Gradient descent memorizes many Gaussians efficiently.
Center manifold analysis can be used in order to investigate the stability of the stationary solutions of various PDEs. This can be done by considering the PDE as an ODE between certain Banach spaces and linearising about the stationary solution. Here we investigate the volume preserving mean curvature flow using such …
In this note, we focus on smooth nonconvex optimization problems that obey: (1) all local minimizers are also global; and (2) around any saddle point or local maximizer, the objective has a negative directional curvature. Concrete applications such as dictionary learning, generalized phase retrieval, and orthogonal ten…
A new method for sparse PCA using orthogonal rotations and soft-thresholding.