New activation functions achieve arbitrary-accuracy Sobolev approximation by fixed-size neural networks.
problem Approximation of Sobolev functions by neural networks
method Elementary Universal Activation Function and Differentiable Universal Activation Functions
result Arbitrary-accuracy Sobolev approximation by fixed-size neural networks
Optimized neural network approximates high-dimensional functions with minimal parameters.
problem Achieving optimal approximation of high-dimensional continuous functions with minimal parameters.
method Developed a neural network with a specific activation function and architecture to achieve super approximation property.
result A composed network with at most 10889d + 10887 nonzero parameters achieves super approximation property, suggesting optimality in parameter growth.
Unified framework proves neural networks' ability to mimic complex tasks.
problem Lack of a single constructive framework for neural network universality.
method Introduces neural network approximate identity (nAI) and proves it leads to universality.
result Any nAI activation function is universal.
The paper proves deep neural networks with analytic activation can approximate any function.
problem Approximating functions with neural networks using analytic activation functions.
method Elementary proofs for real and complex networks, Stone-Weierstrass theorem, Mergelyan's theorem.
result Closure of neural network classes equals space of polynomials for analytic activation.
Study a simplified model of multiverse structure with synchronized timelines.
problem Understanding the complex structure of the local multiverse with multiple universes.
method Time-amalgamated globally hyperbolic model of multiverse as a collection of parallel universes.
result Elementary particles are transcosmic strings with multiple endpoints on parallel universes.
Study of Gödel Universe as Lie group with specific metric.
problem Characterize geodesics in the Gödel Universe.
method Geometric theory of optimal control applied to Lie groups with left-invariant Lorentz metrics.
result No closed timelike or isotropic geodesics in the Gödel Universe.
Introduces minimal surfaces to undergraduates.
problem Teaching minimal surfaces to non-specialist undergraduates.
method Elementary calculus of several variables, accessible to second-year undergraduates.
result Minimal surfaces theory accessible to undergraduates.
We calculate the universal character ring of a class of two-generator, one-relator groups. As an application we give a less technical proof of a result in [LT] on the universal character ring of the (-2,3,2n+1)-pretzel knot. We also give an elementary proof of a result in [Ma] on the character variety of the (-2,3,2n+1…
We study the universal character ring of some families of one-relator groups. As an application, we calculate the universal character ring of two-generator one-relator groups whose relators are palindrome, and, in particular, of the (-2,2m+1,2n+1)-pretzel knot for all integers m and n. For the (-2,3,2n+1)-pretzel knot,…
A new Universal Activation Function improves performance across various machine learning tasks.
problem Achieving near optimal performance in different machine learning tasks.
method Optimization algorithms evolve the UAF's parameters to match the optimal activation function for each task.
result The UAF converges to near optimal performance in classification, quantification, and reinforcement learning tasks.
Universal approximation for ODENet and ResNet with a single activation function.
problem Approximating complex dynamical systems with limited vector fields.
method Examined ODENet and ResNet with vector fields composed of a single activation function and affine mapping.
result ODENet and ResNet with restricted vector fields can uniformly approximate those with general vector fields.
The paper shows neural networks can approximate functions over non-compact domains with non-polynomial activation.
problem Approximating functions over non-compact domains using neural networks.
method Using single-hidden-layer feedforward neural networks with non-polynomial activation functions over non-compact subsets of Euclidean spaces.
result Neural networks can approximate functions in weighted Ck-spaces and weighted Sobolev spaces over unbounded domains. This work shows MLPs can approximate monotonic functions without bounded activations.
problem Optimizing MLPs with monotonic constraints and bounded activations.
method Generalized theoretical results showing MLPs with non-negative weights and saturating activations are universal approximators.
result MLPs with non-negative weights and saturating activations are universal approximators for monotonic functions.
Echo state networks with random weights can approximate any continuous system.
problem Approximating continuous dynamical systems using echo state networks.
method Randomly generated internal weights and a sampling procedure for activation functions.
result Echo state networks with random weights can approximate any continuous casual time-invariant operators with high probability.
Minimum width for ReLU networks to approximate L^p functions is max(d_x+1, d_y).
problem Characterizing the minimum width for ReLU networks to approximate L^p functions.
method Analyzing networks with ReLU activation functions and proving the minimum width required.
result The minimum width required for the universal approximation of L^p functions is exactly max(d_x+1, d_y).
The goal of this article is to give an elementary introduction to Dirac geometry and group-valued moment maps, via pure spinors. The material is based on my lectures at the summer school on 'Poisson geometry in Mathematics and Physics' at Keio University, June 2006.
Complex-valued neural networks can approximate any continuous function with bounded widths and depths.
problem Approximating continuous functions with complex-valued neural networks of bounded widths and depths.
method Analyzing activation functions and proving universality for complex-valued networks.
result Deep narrow complex-valued networks are universal if and only if their activation function is neither holomorphic, nor antiholomorphic, nor R-affine. Complex-valued neural networks can approximate any continuous function.
problem Generalizing the universal approximation theorem to complex-valued networks.
method Characterizing activation functions for complex networks to approximate any continuous function.
result Different activation functions are required for deep vs shallow complex networks to achieve universal approximation.
Active subspace is a model reduction method widely used in the uncertainty quantification community. In this paper, we propose analyzing the internal structure and vulnerability and deep neural networks using active subspace. Firstly, we employ the active subspace to measure the number of "active neurons" at each inter…
MLPs can approximate any function in context, challenging the importance of in-context universality.
problem Understanding why transformers are more effective than classical models.
method Proved MLPs with trainable activation functions are universal in context.
result Transformer success is likely due to factors other than in-context universality.
Consider a finite dimensional (generally reducible) polynomial representation ρof GL_n. A projective compactification of GL_n is the closure of ρ(GL_n) in the space of all operators defined up to a factor (this class of spaces can be characterized as equivariant projective normal compactifications of GL_n). We give an …
Neural networks can approximate functions uniformly across various measures.
problem Universal approximation of functions across different probability measures.
method Proving neural networks are dense in Orlicz spaces, extending classical theorems.
result Neural networks uniformly approximate functions for weakly compact families of measures.
The universal approximation property of various machine learning models is currently only understood on a case-by-case basis, limiting the rapid development of new theoretically justified neural network architectures and blurring our understanding of our current models' potential. This paper works towards overcoming th…
Deep neural networks can interpolate any dataset in the overparametrized regime.
problem Interpolating any dataset with deep neural networks in the overparametrized regime.
method Proving universal approximations and interpolating any dataset with deep neural networks, considering specific conditions on activation functions.
result Interpolation of any dataset is possible in the overparametrized regime with deep neural networks.
We demonstrate that in residual neural networks (ResNets) dynamical isometry is achievable irrespectively of the activation function used. We do that by deriving, with the help of Free Probability and Random Matrix Theories, a universal formula for the spectral density of the input-output Jacobian at initialization, in…
Study proves deep narrow RNNs can approximate any function, with minimum width independent of data length.
problem Proving universality of deep narrow RNNs with bounded widths.
method Analyzing RNNs as dynamical systems, proving universality for deep narrow structures with specific widths.
result Minimum width for universality of deep narrow RNNs is independent of data length.
SympNets identify Hamiltonian systems from data using linear, activation, and gradient modules.
problem Identifying Hamiltonian systems from data.
method Composition of linear, activation, and gradient modules; universal approximation theorems.
result SympNets can approximate arbitrary symplectic maps and generalize well to various Hamiltonian systems.
Dropout schedules can be optimized to significantly reduce model test loss.
problem Improving model performance in neural networks.
method Developed a mean-field theory of dropout at the edge of chaos, proposing front-loaded dropout schedules.
result Front-loaded dropout schedules reduce test loss by 18-35% over constant dropout.
Minimum width for ReLU networks on compact domain is exactly max{d_x, d_y, 2}
problem Characterizing the minimum width for ReLU networks to approximate functions on compact domains
method Analyzing the minimum width for Lp approximation of Lp functions from [0,1]d to Rdy using ReLU-like activation functions result The minimum width for Lp approximation on a compact domain is exactly max{d_x, d_y, 2} for ReLU-like activation functions A Minkowski class is a closed subset of the space of convex bodies in Euclidean space Rn which is closed under Minkowski addition and non-negative dilatations. A convex body in Rn is universal if the expansion of its support function in spherical harmonics contains non-zero harmonics of all orders. If K is universal, t…
We propose the point process model as the Poissonian-like stochastic sequence with slowly diffusing mean rate and adjust the parameters of the model to the empirical data of trading activity for 26 stocks traded on NYSE. The proposed scaled stochastic differential equation provides the universal description of the trad…
o1Neuro neural network approximates complex functions and converges quickly.
problem Approximating complex functions and ensuring convergence in neural networks.
method Sparse indicator activation neurons, population and sample level convergence properties.
result o1Neuro achieves optimal model approximation and convergence with high probability.
The classical Universal Approximation Theorem holds for neural networks of arbitrary width and bounded depth. Here we consider the natural `dual' scenario for networks of bounded width and arbitrary depth. Precisely, let n be the number of inputs neurons, m be the number of output neurons, and let ρ be any nonaff…
Paper proves neural networks can be approximated using interval bounds.
problem Verifying safety and robustness of neural networks.
method Introduces interval universal approximation (IUA) theorem for neural networks.
result Neural networks can be approximated using interval bounds for any continuous function and squashable activation functions.
Residual networks with block width max(d_x, d_y) approximate all functions.
problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.
Proves DCNNs with expansive convolution are strongly universally consistent.
problem Theoretical consistency of deep convolutional neural networks (DCNNs).
method Empirical risk minimization on DCNNs with expansive convolution (with zero-padding).
result DCNNs with expansive convolution are strongly universally consistent.
Mixtures of neural operators reduce active complexity in operator learning.
problem Reduction of active complexity in operator learning models.
method Constructive comparison between routed mixtures of neural operators (MoNOs) and a fixed single-neural-operator construction.
result Every scalar uniformly continuous nonlinear operator can be approximated by a MoNO whose active expert has smaller depth, width, and rank scaling.
This research finds three meta-indicators for university rankings.
problem Complexity in university ranking systems.
method Interpretable machine learning approach.
result Identified three meta-indicators: time, space, and relationships.
The paper proves neural networks with ReLU and softmax can approximate any function.
problem Approximating functions and class labels in neural networks.
method Extended universal approximator theory to neural networks with ReLU and softmax.
result Neural networks with ReLU and softmax can approximate any function and class labels.
Let X be a finite aspherical CW-complex whose fundamental group π1(X) possesses a subnormal series π1(X)⊳Gm⊳...⊳G0 with a non-trivial elementary amenable group G0. We investigate the L2-invariants of the universal covering of such a CW-complex X. We show that the Novikov-Shubin invarian…
We show that deep narrow Boltzmann machines are universal approximators of probability distributions on the activities of their visible units, provided they have sufficiently many hidden layers, each containing the same number of units as the visible layer. We show that, within certain parameter domains, deep Boltzmann…
A new KAN variant uses sinusoidal activations to approximate functions.
problem Approximating multivariable functions using neural networks.
method Replacing inner and outer functions in Kolmogorov-Arnold representation with weighted sinusoidal functions.
result The new KAN variant outperforms fixed-frequency Fourier transform and achieves comparable performance to MLPs.
Deep residual networks can approximate any continuous function using control theory.
problem Universal approximation capabilities of deep residual neural networks.
method Relating residual networks to control systems and using Lie algebraic techniques.
result Deep residual networks with adequately deep layers can approximate any continuous function on a compact set.
Under-parameterized networks can either copy or average teacher weights, leading to universal optimal solutions.
problem Approximating a teacher network with an under-parameterized student network.
method Analyzing shallow neural networks with erf activation function and unitary teacher weights, proving copy-average configurations are critical points and finding the optimal solution.
result The optimal solution for under-parameterized networks has a universal structure, whether copying or averaging teacher neurons.
This paper is devoted to an elementary new construction of 1-singular Gelfand-Tsetlin modules using complex geometry. We introduce a universal ring Do together with the vector space S=S(Do) with basis Bo=B(Do) formed from some local distributi…
Sparse Canonical Correlation Analysis (CCA) has received considerable attention in high-dimensional data analysis to study the relationship between two sets of random variables. However, there has been remarkably little theoretical statistical foundation on sparse CCA in high-dimensional settings despite active methodo…
Given a piecewise linear (PL) function p defined on an open subset of Rn, one may construct by elementary means a unique polyhedron with multiplicities $\D(p)$ in the cotangent bundle Rn×Rn∗ representing the graph of the differential of p. Restricting to dimension 2, we show that any smooth functi…
Global group laws connect equivariant bordism rings to formal group laws.
problem Establishing connections between equivariant bordism rings and formal group laws.
method Global homotopy theory framework; proving isomorphisms and universal properties.
result Equivariant bordism rings are isomorphic to Lazard rings for abelian Lie groups.