A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Wide neural networks converge to Gaussian processes, improving generalization.
problem Understanding the generalization of wide neural networks, especially deep equilibrium models.
method Investigation of deep equilibrium models (DEQs) with infinite-depth layers, focusing on their convergence to Gaussian processes as width and depth approach infinity.
result Wide DEQs converge to Gaussian processes, maintaining generalization performance.
We describe a procedure for creating infinite families of hyperbolic knots having unique minimal genus Seifert surface. A large subset of these knots have the further property that the surface cannot be the sole compact leaf of a depth one foliation of the knot exterior.
We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation is exponential in the ratio of network depth to width. Thus, even in the limit of infinite overparameterization, the NTK is not determinist…
We obtain new invariants of topological link concordance and homology cobordism of 3-manifolds from Hirzebruch-type intersection form defects of towers of iterated p-covers. Our invariants can extract geometric information from an arbitrary depth of the derived series of the fundamental group, and can detect torsion wh…
Recent work by Jacot et al. (2018) has shown that training a neural network using gradient descent in parameter space is related to kernel gradient descent in function space with respect to the Neural Tangent Kernel (NTK). Lee et al. (2019) built on this result by establishing that the output of a neural network traine…
This paper explores how neural network width and depth behave as they approach infinity.
problem Understanding the behavior of neural functions as width and depth go to infinity.
method Formal definition of commutativity framework, study of neural covariance kernel, novel proof techniques.
result Taking width and depth to infinity in a deep neural network with skip connections results in the same covariance structure, regardless of the order of taking limits.
When the parameters are independently and identically distributed (initialized) neural networks exhibit undesirable properties that emerge as the number of layers increases, e.g. a vanishing dependency on the input and a concentration on restrictive families of functions including constant functions. We consider parame…
This article concerns the expressive power of depth in deep feed-forward neural nets with ReLU activations. Specifically, we answer the following question: for a fixed din≥1, what is the minimal width w so that neural nets with ReLU activations, input dimension din, hidden layer widths at most w, and …
The Baumslag-Solitar groups: BS(m,n)=<x,y| x y^{m} x^{-1} = y^{n}> are some of the simplest interesting infinite groups which are not lattices in Lie groups. They have been studied in depth from the point of view of combinatorial group theory. It is natural to ask if the geometric approach to the theory of infinite gro…
A generic geodesic on a finite area, hyperbolic 2-orbifold exhibits an infinite sequence of penetrations into a neighborhood of a cone singularity, so that the sequence of depths of maximal penetration has a limiting distribution. The distribution function is the same for all such surfaces and is described by a fairly …
Large neural networks learn low-dimensional representations that balance complexity and regularity.
problem Understanding the tradeoff between low-dimensional representations and complexity in deep neural networks.
method Computed finite depth corrections to reveal a measure of regularity that bounds the pseudo-determinant of the Jacobian.
result Proved the conjectured bottleneck structure in learned features as network depth increases, showing almost all hidden representations are approximately low-dimensional and weight matrices have singular values close to 1.
This paper investigates the approximation power of three types of random neural networks: (a) infinite width networks, with weights following an arbitrary distribution; (b) finite width networks obtained by subsampling the preceding infinite width networks; (c) finite width networks obtained by starting with standard G…