Shared classical randomness improves quantum generative models' output distributions.
problem Improving generative performance of shallow unitary quantum models.
method Introducing stochasticity into unitary quantum models via shared classical randomness.
result Shared classical randomness allows shallow unitary quantum models to represent a strictly larger family of distributions.
Deeper neural networks can better approximate certain natural functions than shallower ones.
problem Approximating natural functions with neural networks.
method Depth-based separation results for feed-forward neural networks.
result Deeper networks can better approximate certain types of functions than shallower ones.
Unified theorem for deep and shallow joint-equivariant machines.
problem Universal approximation of joint-equivariant machines.
method Constructive universal approximation theorem based on ridgelet transform.
result Unified approximation of deep and shallow networks.
New algorithm proves deep networks can learn better than shallow ones.
problem Understanding the power difference between shallow and deep neural networks.
method Identifying a class of Boolean functions and proving that logarithmic-depth networks can learn them efficiently using hierarchical reconstruction.
result First algorithmic separation between constant-depth and logarithmic-depth neural networks.
Deep CNN architectures improve neonatal seizure detection accuracy.
problem Improving EEG-based neonatal seizure detection accuracy.
method Design and test of deep convolutional networks of varying depths compared to a shallow SVM-based detector.
result A deep 11-layer CNN architecture significantly outperforms shallow architectures, improving AUC90 from 82.6% to 86.8%.
Paper transforms deep rectifier networks into shallow ones for analysis.
problem Understanding the complexity of deep neural networks.
method Transformation of deep rectifier networks into shallow ones.
result Shallow networks can represent deep networks with fewer functions.
This work explores the relation between depth and expressivity in neural networks.
problem Understanding the power of depth in neural networks and its relation to gradient-based optimization.
method Depth separation argument for distributions with fractal structure, proving that deep networks can express fine details efficiently but shallow ones cannot.
result The success of learning deep networks depends on whether the distribution can be well approximated by shallower networks.
Deep nets outperform shallow nets in complex feature realization.
problem Realizing complex data features with deep nets.
method Refined covering number estimates and analysis of approximation rates.
result Deep nets can improve performance without additional capacity costs for complex features.
Estimates neural network error approximating compact sets.
problem Approximating compact subsets from Banach spaces with neural networks.
method Estimates error rates for neural networks of varying width and depth.
result Depth is crucial for better approximation rates, width alone does not improve.
The study analyzes games and social hierarchies, incorporating luck and depth of competition.
problem Analyzing patterns of wins and losses in games and social hierarchies.
method Generalized probabilistic models incorporating luck and depth of competition.
result Social competition tends to be deeper with many distinct levels, but there is often a chance of upset victories.
Study shows shallow ReLU networks struggle with high-dimensional Lipschitz functions.
problem Expressing high-dimensional Lipschitz functions with shallow ReLU networks.
method Established lower bounds on shallow network complexity for polynomial approximation.
result Shallow ReLU networks suffer from the curse of dimensionality for Lipschitz functions.
Deeper networks are better for local labels, but shallower for global labels.
problem Understanding the effect of depth in overparameterized neural networks.
method Introduced local and global labels to investigate the advantage of depth.
result Deeper networks are better for local labels, shallower for global labels.
Deep networks can learn functions approximated by shallow networks, but not all functions.
problem The learnability of functions by deep neural networks and the approximation capacity of simpler classes.
method Study the connection between learnability and approximation capacity of functions by deep neural networks and simpler classes.
result A necessary condition for a function to be learnable by deep neural networks is to be approximable by shallow networks.
Study how depth affects inference in deep Bayesian neural networks.
problem Understanding how depth impacts inference in overparameterized linear Bayesian neural networks.
method Interpreting finite deep linear Bayesian neural networks as scale mixtures of Gaussian process predictors.
result Advances analytical understanding of how depth affects inference in a simple class of Bayesian neural networks.
AutoGrow automatically discovers optimal depth in DNNs.
problem Designing optimal depth in deep neural networks is difficult and time-consuming.
method AutoGrow grows new layers in a seed architecture if it improves accuracy; stops if no improvement. Robust policies generalize to different architectures and datasets.
result AutoGrow discovers near-optimal depth on various datasets, improving accuracy-computation trade-off in ResNets.
New connection between DNNs and Sharkovsky's Theorem for depth-width trade-offs.
problem Understanding why some functions are hard to represent by shallow ReLU networks.
method Connection to Sharkovsky's Theorem and analysis of dynamical systems.
result Lower bounds for width needed to represent periodic functions as a function of depth.
SCAN divides deep neural networks into shallow classifiers for efficient deployment.
problem Explosive growth in storage and computation limits deep neural networks on edge devices.
method SCAN divides networks into shallow classifiers, uses attention modules and knowledge distillation, and employs a threshold-controlled scalable inference mechanism.
result SCAN achieves significant performance gain on CIFAR100 and ImageNet without hyper-parameter adjustments.
CHOOSE enhances shallow Transformers for wireless symbol detection.
problem Improving wireless symbol detection with shallow Transformers.
method Introducing autoregressive latent reasoning steps within hidden space.
result Lightweight Transformers achieve comparable performance to deep models.
Study defends shallow neural networks from data-poisoning attacks.
problem Protecting shallow neural networks from adversarial attacks during training.
method Developed a non-gradient stochastic algorithm for depth-2 neural networks, proving near-optimal trade-offs.
result Demonstrated improved performance over stochastic gradient descent under various data distributions.
Study shows depth improves generalization in deep learning models.
problem Understanding why and when depth improves generalization in deep learning.
method Implementation-agnostic state-transition model to analyze depth and generalization.
result Identifies geometric and semigroup mechanisms that keep entropy contribution saturated or polynomial, clarifying depth's statistical advantage.
Deep networks learn hierarchical functions more efficiently than shallow ones.
problem Understanding the advantage of deep neural networks over shallow models.
method Analytical study of learning dynamics and generalization performance of deep networks compared to shallow ones.
result Deep networks reduce effective dimensionality, enabling learning with fewer samples.
Deep and wide ReLU networks learn data-dependent features even in the lazy training regime.
problem Understanding the behavior of neural networks with finite depth and width.
method Analyzing the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network.
result The NTK has a non-trivial evolution during training, with the mean of its first SGD update being exponential in the ratio of depth to width.
Deep nets improve function approximation and learning in high dimensions.
problem Designing neural networks for rotation-invariant function approximation.
method Developed deep neural networks with multiple hidden layers for radial function approximation.
result Deep nets achieve near-optimal function approximation and learning rates not possible by shallow nets.
Neural networks with depth can't be approximated by shallow ones unless exponentially large.
problem Approximating neural networks with limited depth.
method Proving limitations of shallow neural networks using semi-algebraic gates.
result Neural networks with depth are essential for certain tasks, even for common activation functions.
PathCapsNet improves CapsNet by reducing parameters and enhancing performance.
problem Limitations of CapsNet, including excessive parameters and shallow architecture.
method Introducing a deep parallel multi-path version of CapsNet, incorporating depth, max-pooling, regularization, and new routing techniques.
result Better or comparable results to CapsNet with significantly reduced parameter count.
Deep networks outperform shallow ones in Ising model near criticality.
problem Comparing deep vs shallow neural networks for Ising model.
method Trained deep and shallow Boltzmann machines on Ising system data.
result Accuracy depends only on first hidden layer size, not network depth.
Depth helps neural networks learn simpler solutions incrementally.
problem Understanding why deep neural networks generalize well despite complex architectures.
method Formal definition of incremental learning dynamics, theoretical analysis of depth and initialization effects, experiments with various models.
result Incremental learning dynamics can arise in deeper models, but not in shallow ones, under specific conditions.
Theoretical analysis shows RNNs with various nonlinearities benefit from depth efficiency.
problem Theoretical understanding of RNNs' efficiency is limited.
method Extended analysis to RNNs with Rectified Linear Unit (ReLU) and other nonlinearities.
result Various nonlinear RNNs also benefit from depth efficiency.
Two new methods reduce random forest latency and improve accuracy.
problem High latency and memory demands in deep random forest models.
method DiNo and RanBu convert shallow random forests into efficient predictors.
result RanBu matches or exceeds full-depth random forest accuracy with up to 95% reduction in time.
Enhanced ODT with Feature Concatenation boosts learning efficiency.
problem Insufficient learning efficiency of ODT due to linear projections not being transmitted to child nodes.
method Feature Concatenation ( exttt{FC-ODT}) to transmit linear projections along decision paths.
result Experiments show exttt{FC-ODT} outperforms state-of-the-art decision trees with a limited tree depth.
A method to simplify deep neural networks for specific tasks.
problem Reducing deep neural networks to a smaller size while maintaining functionality.
method Advanced Supervised Principal Component Analysis-based shallowing algorithm.
result The method can reduce network depth without significant performance loss.
GD converges faster to flatter minima than gradient flow in shallow networks.
problem Understanding the dynamics of gradient descent in shallow linear networks.
method Analyzing the convergence rate and solution of gradient descent in depth-2 linear neural networks.
result GD converges linearly to flatter minima than gradient flow, even with large step sizes.
New findings show depth is more important than width in neural networks.
problem Understanding the role of width and depth in neural networks.
method Constructed networks with bounded weights and width at most d+2, showing depth plays a more significant role.
result Depth is more important than width in the expressive power of neural networks.
SRFRN accelerates image super-resolution using shallow residual units.
problem High computational complexity and time in deep learning image super-resolution.
method SRFRN uses a bicubic interpolated low-resolution image and residual representative units (RFR) for faster and more efficient high-resolution image reconstruction.
result SRFRN achieves superior performance and faster execution time compared to existing methods.
DA-LSTM adapts LSTM depth to non-uniform data, improving efficiency.
problem Non-uniform information distribution in sequential data cannot be accurately modeled by traditional LSTM.
method Developed DA-LSTM architecture that dynamically adjusts LSTM depth based on information distribution.
result DA-LSTM reduces computation resource usage and convergence time by 41.78% and 46.01% respectively.
DCTN uses tensor networks for image classification, achieving state-of-the-art results.
problem Improving image classification accuracy with deep neural networks.
method Developed a novel deep convolutional tensor network (DCTN) based on Entangled plaquette states (EPS).
result DCTN achieves state-of-the-art results on MNIST and FashionMNIST but overfits on CIFAR10.
Bayesian linear networks reveal optimal depth and width trade-offs.
problem Understanding how depth, width, and dataset size affect model quality in linear networks.
method Zero noise Bayesian inference with Gaussian weight priors and mean squared error.
result Optimal predictions at infinite depth and maximized Bayesian model evidence at infinite depth.
Study shows critical initialisation not crucial for ReLU networks under dropout limits.
problem Effect of initialisation on training speed and generalisation in ReLU networks.
method Large-scale statistical analysis of over 12,000 trained networks.
result Non-critical initialisations perform similarly to critical initialisations in terms of performance.
Wide neural networks converge to Gaussian processes, improving generalization.
problem Understanding the generalization of wide neural networks, especially deep equilibrium models.
method Investigation of deep equilibrium models (DEQs) with infinite-depth layers, focusing on their convergence to Gaussian processes as width and depth approach infinity.
result Wide DEQs converge to Gaussian processes, maintaining generalization performance.
Optimizes neural networks by removing unnecessary layers, improving performance and speed.
problem Finding the optimal depth of neural networks to improve performance and speed.
method Develops a fast end-to-end method for training lightweight neural networks with multiple classifier heads, allowing the model to determine the importance of each head and choosing a single shallow classifier.
result Significantly reduces the number of parameters and accelerates inference, outperforming many standard pruning methods.
Theory explains why DARTS favors deep architectures over shallow ones.
problem DARTS selects architectures with dominated skip connections, leading to performance degradation.
method Theoretical analysis of operations' effects on network optimization; introduces sparse binary gates and path-depth-wise regularization.
result Theoretical proof that architectures with more skip connections converge faster.
Neural networks can approximate any L^p functions on R^n.
problem Approximating functions on unbounded domains with neural networks.
method Monotone sigmoid, ReLU, ELU, Softplus, LeakyReLU activation functions.
result Shallow neural networks can arbitrarily well approximate L^p functions on R^n.
Neural networks can learn Boolean circuits with local correlation.
problem Learning Boolean circuits with neural networks is computationally hard.
method Observing local correlation between input patterns and target labels, focusing on tree-structured Boolean circuits.
result Local correlation determines the success or failure of optimization in learning Boolean circuits.
AR-GANs learn depth and DoF from unlabeled images using aperture rendering and focus cues.
problem Learning depth and DoF from unlabeled natural images with diverse viewpoints and shapes.
method Aperture rendering and focus cues to learn depth and DoF from unlabeled images.
result AR-GANs effectively learn depth and DoF from various datasets, including flower, bird, and face images.
Deep learning can learn compositional functions more efficiently by breaking them into stages.
problem Understanding why deep learning performs better than shallow models in learning compositional functions.
method Analyzed learnability of compositional target functions using a three-layer fitting model trained with layer-wise spectral estimators.
result Learning compositional functions can be simplified by breaking them into stages, reducing the complexity of the learning problem.
Deep neural networks can grok better than shallow ones, showing multi-stage generalization.
problem Understanding the generalization behavior of deep neural networks.
method Empirical replication and analysis of grokking phenomenon in deep MLP models.
result Deep neural networks exhibit multi-stage generalization, with a secondary surge in test accuracy.
We address representational challenges in normalizing flows, particularly depth and conditioning issues.
problem Challenges in training normalizing flows, including vanishing/exploding gradients and poor conditioning.
method Analyzes representational aspects of depth and conditioning in normalizing flows, proving theoretical bounds and investigating phenomena.
result Proves that shallow affine coupling networks are universal approximators in Wasserstein distance if ill-conditioning is allowed.
Residual networks' depth is mathematically equivalent to expanding an implicit ensemble size.
problem Understanding why deep residual networks are effective.
method Formal analysis of residual networks as ensembles of shallow models.
result Increasing network depth is equivalent to expanding the size of an implicit ensemble, revealing a hierarchical structure.