New approach finds minimum width for deep, narrow MLPs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Minimum width for ReLU networks to approximate L^p functions is max(d_x+1, d_y).
Minimum width for ReLU networks on compact domain is exactly max{d_x, d_y, 2}
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
We prove that among all constant width bodies of revolution, the minimum of the ratio of the volume to the cubed width is attained by the constant width body obtained by rotation of the Reuleaux triangle about an axis of symmetry.
Residual networks with block width max(d_x, d_y) approximate all functions.
We characterize the first min-max width of real projective spaces of any dimension. The width is the minimum area over the Clifford hypersurfaces. We also compute the Morse index of the Clifford hypersurfaces in the complex and quaternionic projective spaces.
Uniform convergence of interpolators proven for Gaussian data.
Study proves deep narrow RNNs can approximate any function, with minimum width independent of data length.
Neural networks with DAGs show linearity as width increases.
Study shows how networks converge to minimum norm solutions with regularization.
To each knot one can associated its knot Floer homology , a finitely generated bigraded abelian group. In general, the nonzero ranks of these homology groups lie on a finite number of slope one lines with respect to the bigrading. The width of the homology is, in essence, the largest horizo…
We focus on estimating \emph{a priori} generalization error of two-layer ReLU neural networks (NNs) trained by mean squared error, which only depends on initial parameters and the target function, through the following research line. We first estimate \emph{a priori} generalization error of finite-width two-layer ReLU …
Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with ( being the number of training samples). In this paper, we prove that, for deep networks, a single layer of width following the inpu…
Mathematical framework for minimum enclosing ball problem.
Surface area and mean width of a cylinder (the convex hull of two parallel disks) in R^3 are computed. It is more difficult to obtain analogous results for a cone (the convex hull of a disk D and a point p). Oblique formulas for mean width, as well as those for mean curvature, are new. Let L denote the unique diameter …
Study on folded ribbon knots and their minimum length.
Wide networks are often believed to have a nice optimization landscape, but what rigorous results can we prove? To understand the benefit of width, it is important to identify the difference between wide and narrow networks. In this work, we prove that from narrow to wide networks, there is a phase transition from havi…
Gradient descent converges to a global minimum in nonlinear ReLU implicit networks with linear width.
In this paper, we analyze the effects of depth and width on the quality of local minima, without strong over-parameterization and simplification assumptions in the literature. Without any simplification assumption, for deep nonlinear neural networks with the squared loss, we theoretically show that the quality of local…
The paper provides examples of keen weakly reducible bridge spheres for links in b-bridge position.
We prove that for an -layer fully-connected linear neural network, if the width of every hidden layer is , where and are the rank and the condition number of the input data, and is the output dimension, then gradient descent with Gaussi…
Paper calculates topological complexity of robot movement in narrow aisles.
DLNs dynamics change with variance, leading to saddle-to-saddle training phases.
The paper improves theoretical bounds on deep neural networks' convergence.
AdaLoss optimizes adaptive learning rates for efficient convergence in various models.
The betting CI outperforms classical methods in constructing confidence intervals for bounded means.
We analyze the loss landscape and expressiveness of practical deep convolutional neural networks (CNNs) with shared weights and max pooling layers. We show that such CNNs produce linearly independent features at a "wide" layer which has more neurons than the number of training samples. This condition holds e.g. for the…
The selection of initial parameter values for gradient-based optimization of deep neural networks is one of the most impactful hyperparameter choices in deep learning systems, affecting both convergence times and model performance. Yet despite significant empirical and theoretical analysis, relatively little has been p…
New NTK bounds show deep networks with minimum over-parameterization can still memorize and optimize.
We study Kauffman's model of folded ribbon knots: knots made of a thin strip of paper folded flat in the plane. The ribbonlength is the length to width ratio of such a ribbon, and it turns out that the way the ribbon is folded influences the ribbonlength. We give an upper bound of for the ribbonlength of $…
We consider learning high-dimensional multi-response linear models with structured parameters. By exploiting the noise correlations among responses, we propose an alternating estimation (AltEst) procedure to estimate the model parameters based on the generalized Dantzig selector. Under suitable sample size and resampli…
We develop a theory of securities price formation and dynamics based on quantum approach and without presuming any similarities with quantum mechanics. Disorder introduced by trading environment leads to probability distribution of returns that is not a smooth curve, but a speckle-pattern fluctuating in both price coor…
We define the Wirtinger width of a knot. Then we prove the Wirtinger width of a knot equals its Gabai width. The algorithmic nature of the Wirtinger width leads to an efficient technique for establishing upper bounds on Gabai width. As an application, we use this technique to calculate the Gabai width of approximately …
In this paper, we prove a conjecture published in 1989 and also partially address an open problem announced at the Conference on Learning Theory (COLT) 2015. With no unrealistic assumption, we first prove the following statements for the squared loss function of deep linear neural networks with any depth and any widths…
Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.
This paper shows how deep neural networks can learn rich, independent features that significantly deviate from initialization.
The isospectral problem for p-widths is solved using Zoll metrics on S^2.
Width trees link link invariants and bridge number.
Lectures on deep learning properties in infinite and large-width networks.
Computed p-widths for hemisphere, first for manifolds with boundary.
Polygon -widths are found via billiard trajectories.
A number of results for C-smooth surfaces of constant width in Euclidean 3-space are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…
Computed p-widths for real projective plane.
Study bounds Urysohn width of manifolds under surgeries.
Adaptive gradient methods like AdaGrad are widely used in optimizing neural networks. Yet, existing convergence guarantees for adaptive gradient methods require either convexity or smoothness, and, in the smooth setting, only guarantee convergence to a stationary point. We propose an adaptive gradient method and show t…
Convex geometry explains optimal neural network parameters.
While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…