A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
The dying ReLU refers to the problem when ReLU neurons become inactive and only output 0 for any input. There are many empirical and heuristic explanations of why ReLU neurons die. However, little is known about its theoretical analysis. In this paper, we rigorously prove that a deep ReLU network will eventually die in…
Recently, neural networks in machine learning use rectified linear units (ReLUs) in early processing layers for better performance. Training these structures sometimes results in "dying ReLU units" with near-zero outputs. We first explore this condition via simulation using the CIFAR-10 dataset and variants of two popu…
Activation functions play a key role in providing remarkable performance in deep neural networks, and the rectified linear unit (ReLU) is one of the most widely used activation functions. Various new activation functions and improvements on ReLU have been proposed, but each carry performance drawbacks. In this paper, w…
Kernel sparsity ("dying ReLUs") and lack of diversity are commonly observed in CNN kernels, which decreases model capacity. Drawing inspiration from information theory and wireless communications, we demonstrate the intersection of coding theory and deep learning through the Grassmannian subspace packing problem in CNN…
In this note, we study the integral of the 1-form logxydy−logyxdx over certain plane curves defined by A-polynomials of knots. It is quite surprising that a Chern-Simons type invariant of 3-manifolds, which can be geometrically computed, may be used to get the exact values of those integrals. Th…
We study a variant of decision-theoretic online learning in which the set of experts that are available to Learner can shrink over time. This is a restricted version of the well-studied sleeping experts problem, itself a generalization of the fundamental game of prediction with expert advice. Similar to many works in t…
We prove that, starting at an initial metric g(0)=e2u0(dx2+dy2) on R2 with bounded scalar curvature and bounded u0, the Ricci flow ∂tg(t)=−Rg(t)g(t) converges to a flat metric on R2.
We explicitely compute the essential spectrum of the Laplace-Beltrami operator for p-forms for the class of warped product metrics dσ2=y2ady2+y2bdθ∂M2, where y is a boundary defining function on a compact manifold with boundary M.
The isotropic 3-space \mathbb{I}^{3} is a real affine 3-space endowed with the metric dx^{2}+dy^{2}. In this paper we describe Weingarten and linear Weingarten affine translation surfaces in \mathbb{I}^{3}. Further we classify the affine translation surfaces in \mathbb{I}^{3} that satisfy certain equations in terms of …
A semi-isotropic space is a real affine 3-space endowed with the non-degenerate metric dx^{2}-dy^{2}. The main purpose of this paper is to describe the surfaces of revolution in the semi-isotropic space that satisfy some equations in terms of the position vector and the Laplace operators with respect to the first and t…
This study connects ReLU neural networks to toric geometry to analyze function realization.
problem Determining which continuous piecewise linear functions can be realized by ReLU neural networks.
method Established a connection between toric geometry and ReLU neural networks, defining key structures like the ReLU fan, toric variety, and Cartier divisor.
result Proved a criterion for functions realizable by unbiased shallow ReLU networks using intersection numbers.
We introduce a new Self-Organized Criticality (SOC) model for simulating price evolution in an artificial financial market, based on a multilayer network of traders. The model also implements, in a quite realistic way with respect to previous studies, the order book dy- namics, by considering two assets with variable f…
In this paper we characterize the degenerate elliptic equations F(D^2u)=0 whose viscosity subsolutions, (F(D^2u) \geq 0), satisfy the strong maximum principle. We introduce an easily computed function f(t) for t > 0, determined by F, and we show that the strong maximum principle holds depending on whether the integral …
The paper analyzes deep ReLU CNNs' approximation properties in 2D space.
problem Establishing L2 approximation properties for deep ReLU CNNs.
method Analysis based on decomposition theorem for convolutional kernels, properties of ReLU activation, and connections with one-hidden-layer ReLU NNs.
result Universal approximation theorem for deep ReLU CNNs with classic structure.
Proves existence of optimal shallow neural networks with ReLU activation.
problem Proving the existence of optimal shallow feedforward networks with ReLU activation.
method Proves existence of global minima in the loss landscape for continuous target functions using shallow feedforward neural networks with ReLU activation.
result Existence of global minima in the loss landscape for shallow feedforward networks with ReLU activation.
We present an approach to cohomological dimension theory based on infinite symmetric products and on the general theory of dimension called the extension dimension. The notion of the extension dimension $\ExD(X)$ was introduced by A.N.Dranishnikov \cite {D5} in the context of compact spaces and CW complexes. This pa…
This article concerns the expressive power of depth in neural nets with ReLU activations and bounded width. We are particularly interested in the following questions: what is the minimal width wmin(d) so that ReLU nets of width wmin(d) (and arbitrary depth) can approximate any continuous functio…
Let g=e2u(dx2+dy2) be a conformal metric defined on the unit disk of C. We give an estimate of ∥∇u∥L2,∞(D21) when ∥K(g)∥L1 is small and πr2μ(Brg(z),g)<Λ for any r and z∈D43. Then we will use this estimate to study the Gromov-Hausdor…
We consider the problem of computing the best-fitting ReLU with respect to square-loss on a training set when the examples have been drawn according to a spherical Gaussian distribution (the labels can be arbitrary). Let opt<1 be the population loss of the best-fitting ReLU. We prove: 1. Finding a ReLU wit…
Despite their prevalence in neural networks we still lack a thorough theoretical characterization of ReLU layers. This paper aims to further our understanding of ReLU layers by studying how the activation function ReLU interacts with the linear component of the layer and what role this interaction plays in the success …
We study the approximation properties of random ReLU features through their reproducing kernel Hilbert space (RKHS). We first prove a universality theorem for the RKHS induced by random features whose feature maps are of the form of nodes in neural networks. The universality result implies that the random ReLU features…
The paper explores how ReLU DNNs can represent MPC policies and vice versa.
problem Representing MPC policies as ReLU DNNs and vice versa.
method Developed an approximate method for identifying input-space in ReLU nets resulting in PWA functions over polyhedral regions. Studied inverse multiparametric linear or quadratic programs for reconstruction of constraints and cost functions given a PWA function.
result Identification and representation of MPC policies as ReLU DNNs and vice versa.
Using the large deviation principle (LDP) for a re-scaled fractional Brownian motion BtH where the rate function is defined via the reproducing kernel Hilbert space, we compute small-time asymptotics for a correlated fractional stochastic volatility model of the form $dS_t=S_tσ(Y_t) (\barρ dW_t +ρdB_t), \,dY_t=dB^H…