Batch normalization makes deep residual networks train faster.
problem Training deep residual networks with large depths.
method Downscaling the residual branch by a normalizing factor early in training.
result Normalized residual blocks compute functions close to the identity function early in training.
Deep residual networks trained with gradient descent have small generalization gap.
problem Limited theoretical understanding of why residual networks generalize well.
method Analyzing overparameterized deep residual networks trained by gradient descent.
result Demonstrates that residual networks have a small generalization gap between training and test error.
Generalization bounds derived for neural ODEs and deep residual networks.
problem Understanding the generalization capability of neural ODEs and deep residual networks.
method Lipschitz-based argument and analogy with deep residual networks.
result A generalization bound involving the magnitude of weight matrix differences.
ResGCN detects anomalies in attributed networks by capturing sparsity and nonlinearity.
problem Detecting anomalous nodes in attributed networks.
method Attention-based deep residual modeling using Graph Convolutional Networks.
result ResGCN effectively detects anomalies in attributed networks.
Gradient descent converges globally in deep linear residual networks with ZAS initialization.
problem Optimizing deep linear residual networks for convergence.
method Zero-asymmetric (ZAS) initialization for gradient descent.
result Gradient descent converges to an ε-optimal point in O(L^3 log(1/ε)) iterations.
Residual networks maintain stability, allowing deep learning without degradation.
problem Training error increases with network depth.
method Lyapunov stability analysis and convex decomposition of functions.
result Residual networks maintain stable equilibria, preventing overfitting.
Paper transforms deep rectifier networks into shallow ones for analysis.
problem Understanding the complexity of deep neural networks.
method Transformation of deep rectifier networks into shallow ones.
result Shallow networks can represent deep networks with fewer functions.
Study shows how deep residual networks can be analyzed as shallow network ensembles for optimization.
problem Understanding why deep neural networks can be trained to zero loss despite non-convex optimization landscapes.
method Mean-field analysis of deep residual networks, focusing on their continuum limit as a two-layer network.
result Derives the first global convergence result for multilayer neural networks in the mean-field regime.
Deep residual networks implicitly converge to neural ODEs.
problem Link between discrete and continuous deep learning models.
method Establishing implicit regularization for residual networks towards neural ODEs.
result Deep residual networks initialized as discretizations of neural ODEs converge to such ODEs during training.
Deep residual networks improve short-term power load forecasting.
problem Short-term power load forecasting accuracy and generalization.
method Modified deep residual network with two-stage ensemble strategy and Monte Carlo dropout.
result The proposed model provides accurate load forecasting results and high generalization capability.
Identity parameterization simplifies deep learning models and improves performance.
problem Designing deep neural networks with stable and expressive architectures.
method Theoretical analysis and empirical validation of residual networks with identity parameterization.
result Residual networks with identity parameterization have no spurious local optima and universal expressivity.
Residuals improve deep neural networks without increasing hypothesis complexity.
problem Understanding how residual connections affect hypothesis complexity and generalization.
method Analyzing the covering number of the hypothesis space and deriving a margin-based generalization bound.
result Residual connections do not increase the hypothesis complexity of neural networks.
Deep residual networks can approximate any continuous function using control theory.
problem Universal approximation capabilities of deep residual neural networks.
method Relating residual networks to control systems and using Lie algebraic techniques.
result Deep residual networks with adequately deep layers can approximate any continuous function on a compact set.
Adaptive weights improve physics-informed neural networks and deep operator networks.
problem Training physics-informed neural networks and deep operator networks can be challenging, leading to unsatisfactory accuracy and efficiency.
method Proposes a pointwise adaptive weighting method that balances the residual decay rate across different training points.
result Our proposed approach of balanced residual decay rates offers advantages including bounded weights, high prediction accuracy, fast convergence rate, low training uncertainty, low computational cost, and ease of hyperparameter tuning.
ResNets are effective due to their insensitivity to weight initialization.
problem Training deep neural networks with vanishing/exploding gradients.
method Simplified analysis of ResNet models, demonstrating the effect of batch normalization.
result ResNets are insensitive to weight initialization, improving training stability.
A new deep learning model improves robustness and efficiency in predicting continuous variables.
problem Limited applicability of deep learning in domains with small sample sizes.
method Autoencoder-based residual deep network with shortcut connections.
result Achieves cutting-edge accuracy and efficiency in multiple datasets.
Gradient descent efficiently finds global minima in deep neural networks.
problem Training deep neural networks efficiently and reliably.
method Gradient descent, leveraging the stability of the Gram matrix induced by the network architecture.
result Gradient descent achieves zero training loss in polynomial time for deep over-parameterized neural networks with residual connections.
Deep learning networks are approximated using dynamical systems theory.
problem Understanding the approximation capabilities of deep learning networks.
method Modeling deep residual networks as continuous-time dynamical systems and using approximation theories in Lp. result Established general sufficient conditions for universal approximation of deep residual networks.
Deep learning removes aliasing in MRI scans with fast computation.
problem High computational costs in MR scan reconstruction.
method Deep residual learning networks for magnitude and phase networks.
result Deep learning successfully removes aliasing artifacts from MRI scans.
SRFRN accelerates image super-resolution using shallow residual units.
problem High computational complexity and time in deep learning image super-resolution.
method SRFRN uses a bicubic interpolated low-resolution image and residual representative units (RFR) for faster and more efficient high-resolution image reconstruction.
result SRFRN achieves superior performance and faster execution time compared to existing methods.
A new network learns market conditions and predicts stock performance.
problem Optimizing stock portfolio performance in the US equities market.
method Residual Switching Network combining two ResNets: a switching module and a main module.
result The residual switching network strategy outperformed other models with an average annual Sharpe ratio of 2.22.
Speeds up deep neural networks training by 10x using GPU concurrency.
problem Training deep residual neural networks efficiently.
method Layer-wise parallel training with GPU concurrency and Nonlinear Multigrid.
result 10.2x speedup over traditional techniques.
Proposes RMN for learning long-term dependencies in feed-forward networks.
problem Complicated training of deep RNN architectures.
method Residual Memory Neural Network (RMN) with residual and time-delayed connections.
result RMN and BRMN outperform LSTM and BLSTM networks in learning long-term and hierarchical information.
RDL-Net improves speech enhancement with fewer parameters and better performance.
problem Improving speech enhancement with fewer parameters and better performance.
method Proposes RDL-Net, a CNN combining residual and dense aggregations without over-allocating parameters.
result RDL-Net achieves higher speech enhancement performance with fewer parameters and lower computational requirements.
Theoretical limits of deep residual networks show consistent covariance structures.
problem Understanding the limits of deep residual networks.
method Analyzing the behavior of deep residual networks with skip connections as width and depth approach infinity.
result Theoretical analysis confirms that the covariance structure remains consistent regardless of the order of width and depth.
Paper presents a new method for efficient deep learning with over-complete dictionaries.
problem Learning an over-complete basis for optimal reconstruction without optimization.
method Multiscale Residual Mixture of PCA with a hierarchical approach.
result Exponential decrease of error with depth in a recursive deep approach.
Deep learning enhances options hedging performance.
problem Improving delta hedging for options using neural networks.
method Learning residuals between hedging function and implied Black-Scholes delta using neural networks.
result Deep learning significantly improves hedging performance, often by more than 100%.
New measure assesses deep neural networks' robustness to adversarial attacks.
problem Deep learning's fragility to adversarial attacks limits its adoption in mission-critical applications.
method Introduces residual error as a new performance measure for assessing adversarial robustness.
result Demonstrates effectiveness of residual error in assessing robustness of deep neural networks.
Graph neural networks suffer from oversmoothing, but adding residual connections helps.
problem Oversmoothing in deep graph neural networks where features become indistinguishable.
method Analyzed asymptotic oversmoothing rates with and without residual connections using the multiplicative ergodic theorem.
result Adding residual connections effectively mitigates or prevents oversmoothing.
DeepcomplexMRI uses deep residual networks for faster MRI imaging.
problem Accelerating parallel MR imaging with high accuracy.
method Deep complex convolutional neural network with residual connections and k-space consistency.
result The method can accurately reconstruct multi-channel MRI images.
This paper explains why ResNets generalize better than FFNets using neural tangent kernels.
problem Understanding why deep ResNets generalize better than deep FFNets.
method Using neural tangent kernels to compare the learnability of functions induced by the kernels of ResNets and FFNets.
result The kernel of ResNets does not exhibit degeneracy as depth increases, unlike FFNets.
CrescendoNet is a simple deep CNN outperforming others on benchmark datasets.
problem Improving performance of deep neural networks without residual connections.
method Stacking simple blocks with independent convolution paths, increasing depth linearly.
result CrescendoNet with 15 layers and 4.1M parameters outperforms DenseNet-BC with 250 layers and 15.3M parameters.
Fixup replaces normalization in deep networks, achieving similar stability and performance.
problem The effectiveness of normalization layers in deep neural networks.
method Fixed-update initialization (Fixup) to solve exploding and vanishing gradient problems.
result Residual networks trained with Fixup achieve state-of-the-art performance without normalization.
Improved low-dose CT images with deep learning and framelet denoising.
problem Low-dose X-ray CT images often lack texture and detail.
method Proposed a wavelet residual network combining deep learning and framelet denoising.
result Significantly improved performance in preserving image detail.
This paper improves speech recognition models using deep neural networks.
problem Improving automatic speech recognition models.
method Deployment of residual, densely-connected, and highway networks into speech recognition.
result Significant improvements in speech recognition models.
Gradient descent converges to global minima for ResNets with linearly scaled width.
problem Understanding the convergence of deep residual networks with varying network width and dataset size.
method Analyzing the Jacobian of ResNets and applying gradient descent for quadratic loss.
result Gradient descent converges to global minima for ResNets with linearly scaled width and independent of depth.
RED-SC improves deep subspace clustering efficiency.
problem Slow convergence in training deep subspace clustering models.
method Residual Encoder-Decoder network with skip-layer connections and self-expressive layer.
result Training converges much faster with RED-SC.
Stacking improves deep neural network training efficiency.
problem Improving the efficiency of training deep neural networks.
method Proposes stacking as a form of accelerated gradient descent.
result Proves stacking provides accelerated training for certain deep linear residual networks.
New research shows alternative linear connections can outperform identity shortcuts in deep networks.
problem Explaining the effectiveness of shortcut connections in deep neural networks.
method Used variations of the standard residual block with different types of linear connections to build image classification networks.
result Alternative linear connections can be more effective than identity shortcuts in deep networks.
Review of neural network expressivity and architectures.
problem Understanding neural network expressivity across different architectures.
method Comprehensive overview of approximation results for various neural network types.
result Deep neural networks offer advantages over shallow ones for specific function classes.
Develops reversible neural networks for deep learning, improving efficiency and performance.
problem Training and deploying deep neural networks with limited resources and memory.
method Interprets deep residual networks as ODEs, derives reversible architectures, and analyzes stability.
result Reversible neural networks can be trained deeper with less memory and computational resources.
Enhances anomaly detection in high dimensions with pretrained networks.
problem Difficult to characterize anomaly in high-dimensional data.
method Residual adaptation to adjust pretrained networks for anomaly detection.
result Significantly outperforms existing methods on anomaly detection benchmarks.
ResNets promote smoother interpolations than MLPs, enhancing generalization.
problem Understanding the difference in smoothness between ResNets and MLPs.
method Neural Tangent Kernel (NTK) analysis during gradient descent training.
result ResNet's NTK results in smoother interpolations than MLPs.
Effective Gram matrix predicts deep network generalization.
problem Understanding and predicting deep network generalization.
method Derived a differential equation governing generalization gap, analyzed with effective Gram matrix.
result Effective Gram matrix accurately predicts test loss during training.
RFRBoost uses random features to boost deep residual neural networks, improving performance and computational efficiency.
problem Improving performance of deep residual neural networks (RFNNs) while preserving convex optimization benefits.
method Random Feature Representation Boosting (RFRBoost) using boosting theory and random features at each layer.
result RFRBoost significantly outperforms RFNNs and end-to-end trained MLP ResNets in small- to medium-scale tabular datasets.
DAS-PINNs uses deep learning to solve complex PDEs more accurately.
problem Solving high-dimensional PDEs with high accuracy.
method Deep neural networks and generative models for adaptive sampling.
result DAS-PINNs significantly improves solution accuracy for low regularity and high-dimensional problems.
New deep architecture improves stability and robustness of neural networks.
problem Improving stability and robustness of neural networks.
method Proposes a new deep architecture using implicit residual blocks with fixed point definitions.
result Improved stability and generalization power with fewer hyperparameters.
New method uses free probability to improve deep ResNet initialization speed.
problem Improving initialization speed of deep ResNets.
method Introducing a novel analytical tool in free probability for non-Hermitian random matrices.
result Proposed initialization scheme learns at a speed of orders of magnitudes faster.