Gradient descent methods for deep ReLU networks achieve optimal generalization rates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generalization bounds derived for neural ODEs and deep residual networks.
The paper analyzes how low-rank layers in neural networks improve generalization.
Improved deep neural network generalization through noise resilience.
Taking inspiration from biological evolution, we explore the idea of "Can deep neural networks evolve naturally over successive generations into highly efficient deep neural networks?" by introducing the notion of synthesizing new highly efficient, yet powerful deep neural networks over successive generations via an ev…
Deep ReLU networks can be simplified to a three-layer model.
Deep neural networks' infinite-width behavior approximated by Gaussian models.
It is well established that neural networks with deep architectures perform better than shallow networks for many tasks in machine learning. In statistical physics, while there has been recent interest in representing physical data with generative modelling, the focus has been on shallow neural networks. A natural ques…
Novel framework explains generalization in deep neural networks.
Bayesian methods enhance deep learning models by improving reliability and uncertainty.
With the growth of deep learning, how to describe deep neural networks unifiedly is becoming an important issue. We first formalize neural networks mathematically with their directed graph representations, and prove a generation theorem about the induced networks of connected directed acyclic graphs. Then, we set up a …
The paper explores neural scaling laws for deep operator networks, offering a theoretical foundation.
Advancements in deep generative models such as generative adversarial networks and variational autoencoders have resulted in the ability to generate realistic images that are visually indistinguishable from real images, which raises concerns about their potential malicious usage. In this paper, we present an analysis o…
Deep networks become equivalent to linear models in large data regimes.
Deep neural networks can grok better than shallow ones, showing multi-stage generalization.
Deep learning has been widely applied and brought breakthroughs in speech recognition, computer vision, and many other domains. The involved deep neural network architectures and computational issues have been well studied in machine learning. But there lacks a theoretical foundation for understanding the approximation…
Deep neural networks and in particular, deep neural classifiers have become an integral part of many modern applications. Despite their practical success, we still have limited knowledge of how they work and the demand for such an understanding is evergrowing. In this regard, one crucial aspect of deep neural network c…
Deep residual networks (ResNets) have demonstrated better generalization performance than deep feedforward networks (FFNets). However, the theory behind such a phenomenon is still largely unknown. This paper studies this fundamental problem in deep learning from a so-called "neural tangent kernel" perspective. Specific…
New bound for neural nets on non-iid data.
Geodesics found in deep linear networks.
The skip-connections used in residual networks have become a standard architecture choice in deep learning due to the increased training stability and generalization performance with this architecture, although there has been limited theoretical understanding for this improvement. In this work, we analyze overparameter…
This work analyzes how different layers in deep neural networks contribute to generalization error.
Deep neural networks with adversarial training achieve sup-norm convergence for nonparametric regression.
A recent line of research on deep learning focuses on the extremely over-parameterized setting, and shows that when the network width is larger than a high degree polynomial of the training sample size and the inverse of the target error , deep neural networks learned by (stochastic) gradient descent enjoy …
New kernel connects deep learning to optimization, improving generalization.
Characterizes deep neural network weight space for adversarial attacks.
Analyzes minima of deep linear networks with weight decay.
Lecture notes on linear neural networks for deep learning optimization and generalization.
Neuromorphic hardware tends to pose limits on the connectivity of deep networks that one can run on them. But also generic hardware and software implementations of deep learning run more efficiently for sparse networks. Several methods exist for pruning connections of a neural network after it was trained without conne…
Study compares random and learned features in deep Bayesian linear models.
We combine Riemannian geometry with the mean field theory of high dimensional chaos to study the nature of signal propagation in generic, deep neural networks with random weights. Our results reveal an order-to-chaos expressivity phase transition, with networks in the chaotic phase computing nonlinear functions whose g…
The study examines deep convolutional neural networks and their learning ability.
Deep neural networks, in particular convolutional neural networks, have become highly effective tools for compressing images and solving inverse problems including denoising, inpainting, and reconstruction from few and noisy measurements. This success can be attributed in part to their ability to represent and generate…
DAMNETS generates complex network dynamics models.
The generalization error of deep neural networks via their classification margin is studied in this work. Our approach is based on the Jacobian matrix of a deep neural network and can be applied to networks with arbitrary non-linearities and pooling layers, and to networks with different architectures such as feed forw…
The paper develops generalization bounds for deep compound Gaussian neural networks.
Deep networks generalize well due to hidden mechanisms like renormalization.
Deep neural networks are widely used in various domains. However, the nature of computations at each layer of the deep networks is far from being well understood. Increasing the interpretability of deep neural networks is thus important. Here, we construct a mean-field framework to understand how compact representation…
We present a comprehensive study of multilayer neural networks with binary activation, relying on the PAC-Bayesian theory. Our contributions are twofold: (i) we develop an end-to-end framework to train a binary activated deep neural network, (ii) we provide nonvacuous PAC-Bayesian generalization bounds for binary activ…
This paper develops a novel deep recurrent neural network for sequential signal reconstruction.
The paper explores stability and generalization of deep GCNs.
Deep learning architectures have proved versatile in a number of drug discovery applications, including the modelling of in vitro compound activity. While controlling for prediction confidence is essential to increase the trust, interpretability and usefulness of virtual screening models in drug discovery, techniques t…
Deep neural networks can generate any 2D distribution with high accuracy.
Deep neural networks have introduced novel and useful tools to the machine learning community. Other types of classifiers can potentially make use of these tools as well to improve their performance and generality. This paper reviews the current state of the art for deep learning classifier technologies that are being …
The paper develops a deep neural network estimator for weakly dependent processes with various loss functions.
Deep neural networks classify unbounded Gaussian mixture data without dimensionality issues.
Deep neural network with l_1-regularization achieves nearly optimal risk bounds.
There has been a growing interest in expressivity of deep neural networks. However, most of the existing work about this topic focuses only on the specific activation function such as ReLU or sigmoid. In this paper, we investigate the approximation ability of deep neural networks with a broad class of activation functi…