A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We present a new, unifying approach following some recent developments on the complexity of neural networks with piecewise linear activations. We treat neural network layers with piecewise linear activations as tropical polynomials, which generalize polynomials in the so-called (max,+) or tropical algebra, with pos…
Paper proposes LANN to measure model complexity of neural networks with curve activation functions.
problem Measuring model complexity of neural networks with curve activation functions.
method Proposes LANN, a piecewise linear framework to approximate curve activation functions, and derives complexity measure based on the number of linear regions.
result Demonstrates positive correlation between overfitting and model complexity during training.
Artificial neural networks typically have a fixed, non-linear activation function at each neuron. We have designed a novel form of piecewise linear activation function that is learned independently for each neuron using gradient descent. With this adaptive activation function, we are able to improve upon deep neural ne…
We study the complexity of functions computable by deep feedforward neural networks with piecewise linear activations in terms of the symmetries and the number of linear regions that they have. Deep networks are able to sequentially map portions of each layer's input-space to the same output. In this way, deep models c…
The approximation power of general feedforward neural networks with piecewise linear activation functions is investigated. First, lower bounds on the size of a network are established in terms of the approximation error and network depth and width. These bounds improve upon state-of-the-art bounds for certain classes o…
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large…
The developments of deep neural networks (DNN) in recent years have ushered a brand new era of artificial intelligence. DNNs are proved to be excellent in solving very complex problems, e.g., visual recognition and text understanding, to the extent of competing with or even surpassing people. Despite inspiring and enco…
We introduce a variational framework to learn the activation functions of deep neural networks. Our aim is to increase the capacity of the network while controlling an upper-bound of the actual Lipschitz constant of the input-output relation. To that end, we first establish a global bound for the Lipschitz constant of …
This article provides an attempt to extend concepts from the theory of Riemannian manifolds to piecewise linear spaces. In particular we propose an analogue of the Ricci tensor, which we give the name of an Einstein vector field. On a given set of piecewise linear spaces we define and discuss (normalized) Ricci flows. …
We prove that every piecewise linear manifold of dimension up to four on which a finite group acts by piecewise linear homeomorphisms admits a compatible smooth structure with respect to which the group acts smoothly. This solves a challenge posed by Thurston in dimension three and confirms a conjecture by Kwasik and L…
In this paper, we introduce a bordism category CdPL whose objects are bundles of closed (d−1)-dimensional piecewise linear manifolds and whose morphisms are bundles of d-dimensional piecewise linear cobordisms. In the main theorem of this article, we show that the classifying space $B\mathcal{C}_d^{…
We propose to optimize the activation functions of a deep neural network by adding a corresponding functional regularization to the cost function. We justify the use of a second-order total-variation criterion. This allows us to derive a general representer theorem for deep neural networks that makes a direct connectio…
We consider smooth isotropic immersions from the 2-dimensional torus into R2n, for n≥2. When n=2 the image of such map is an immersed Lagrangian torus of R4. We prove that such isotropic immersions can be approximated by arbitrarily C0-close piecewise linear isotropic maps. If n≥3 the piece…
It is well-known that the expressivity of a neural network depends on its architecture, with deeper networks expressing more complex functions. In the case of networks that compute piecewise linear functions, such as those with ReLU activation, the number of distinct linear regions is a natural measure of expressivity.…
Hilbert initiated the standpoint in foundations of mathematics. From this standpoint, we allow only a finite number of repetitions of elementary operations when we construct objects and morphisms. When we start from a subset of a Euclidean space. Then we assume that any element of the line has only a finite number of c…