Study of infinitely deep but narrow neural networks using NTK theory.
problem Analyzing the role of depth in deep learning with overparameterized networks.
method Infinite-depth limit analysis of MLP and CNN using Neural Tangent Kernel (NTK) theory.
result Established trainability guarantee for infinitely deep but narrow neural networks.
We show that deep narrow Boltzmann machines are universal approximators of probability distributions on the activities of their visible units, provided they have sufficiently many hidden layers, each containing the same number of units as the visible layer. We show that, within certain parameter domains, deep Boltzmann…
Deep and narrow neural networks can collapse to incorrect states.
problem The collapse of deep and narrow neural networks to incorrect states.
method Numerical and theoretical analysis of deep and narrow neural networks with ReLU activation.
result Deep and narrow neural networks can converge to erroneous mean or median states with high probability.
New approach finds minimum width for deep, narrow MLPs.
problem Finding the minimum width for deep, narrow MLPs to approximate continuous functions.
method Proposes a framework to simplify finding minimum width into determining a geometrical function w(dx,dy) based on input and output dimensions. result Proves that w(dx,dy) equals the optimal minimum width for deep, narrow MLPs to achieve universality. Narrow neural networks have unbounded decision regions.
problem Understanding decision regions of narrow neural networks.
method Analyzing decision regions of neural networks with width ≤ input dimension.
result All connected components of decision regions are unbounded.
Deep narrow networks can approximate any continuous function.
problem Approximating continuous functions with neural networks of bounded width and arbitrary depth.
method Showed neural networks of arbitrary depth, width n+m+2, and activation function ρ is dense in C(K;Rm) for K⊆Rn with K compact. result Neural networks of bounded width and arbitrary depth can approximate any continuous function.
Study proves deep narrow RNNs can approximate any function, with minimum width independent of data length.
problem Proving universality of deep narrow RNNs with bounded widths.
method Analyzing RNNs as dynamical systems, proving universality for deep narrow structures with specific widths.
result Minimum width for universality of deep narrow RNNs is independent of data length.
Embedding principle explains loss landscape of deep neural networks.
problem Understanding the structure of loss landscapes in deep neural networks.
method Proposed an embedding principle that critical points of narrower DNNs can be embedded to critical points of wider DNNs.
result Wide DNNs are often attracted by highly-degenerate critical points embedded from narrower DNNs.
Improved bounds on neural network expressivity.
problem Understanding neural network expressivity and approximation capabilities.
method Improved bounds on the maximal number of linear regions of ReLU-networks.
result New insights into the expressivity of neural networks.
Deep learning uncovers patterns between knot types.
problem Discovering connections between combinatorial and hyperbolic knot invariants.
method Statistical approach using linear regression and deep learning.
result Revealed empirical connections between knot types.
Custom narrow-precision representations boost DNN inference speed by 7.6x with minimal accuracy loss.
problem Improving computational efficiency of deep neural networks.
method Exploring and utilizing unconventional narrow-precision floating-point representations for DNN weights and activations.
result Average speedup of 7.6x with less than 1% accuracy loss.
Study compares random and learned features in deep Bayesian linear models.
problem Understanding how feature learning affects generalization in deep learning.
method Comparing deep random feature models to deep networks with trained layers.
result Random feature models can display double-descent behavior, while deep networks do not.
We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Montúfar and Ay,…
We review recent results about the maximal values of the Kullback-Leibler information divergence from statistical models defined by neural networks, including naive Bayes models, restricted Boltzmann machines, deep belief networks, and various classes of exponential families. We illustrate approaches to compute the max…
Investigates the impact of narrow banking on macroeconomics.
problem The risks and benefits of a full reserve requirement on demand deposits.
method Extended Goodwin-Keen model with time deposits and central bank reserves; numerical examples.
result Narrow banking does not reduce economic growth but improves financial stability.
Wide neural networks with narrow bottlenecks behave like deep Gaussian processes.
problem Understanding the behavior of neural networks with narrow layers in the wide limit.
method Analyzing the wide limit of BNNs with narrow bottlenecks, showing they behave like a composition of GPs.
result Wide neural networks with narrow bottlenecks form a composition of GPs, termed a bottleneck NNGP.
Statistical physics explains deep learning's feature learning capacity.
problem Understanding neural networks' ability to learn complex features.
method Study of a multi-layer perceptron in the interpolation regime.
result Optimal learning requires specialization across layers and neurons.
Deep ResNets with single-neuron hidden layers can approximate any function.
problem The challenge of universal approximation by deep neural networks.
method A ResNet architecture with one neuron per hidden layer in each module.
result ResNet with one-neuron hidden layers is a universal approximator.
Study predicts risk of true-lumen narrowing after ATAAD surgery using CT data.
problem Early post-surgery risk assessment for aortic dissection patients.
method Retrospective study with CT data, derived cross-sectional shapes, form factor (FF) for morphology assessment, linear discriminant analysis (LDA) for risk classification, LOPO-CV for prediction.
result Machine-learning model accurately predicts risk for all high-risk patients and low-risk patients, potentially reducing hospital visits.
Study on ion travel time on curved surfaces.
problem Mean first passage time of ion on curved surfaces.
method Layer potential argument and microlocal analysis.
result Derivation of mean first passage time and spatial average.
Gradient descent finds global min in wide deep linear networks.
problem Optimizing deep linear neural networks with limited width.
method Proved convergence rate for gradient descent with wide layers.
result Gradient descent converges linearly to global min in wide layers.
This paper improves deep neural network approximation for fully connected networks, achieving optimal convergence rates.
problem Improving approximation of fully connected deep neural networks for optimal convergence rates.
method Deriving approximation bounds specifically for a narrower fully connected deep neural network.
result Achieves an optimal rate (up to a logarithmic factor) for fully connected deep neural networks.
Complex-valued neural networks can approximate any continuous function with bounded widths and depths.
problem Approximating continuous functions with complex-valued neural networks of bounded widths and depths.
method Analyzing activation functions and proving universality for complex-valued networks.
result Deep narrow complex-valued networks are universal if and only if their activation function is neither holomorphic, nor antiholomorphic, nor R-affine. Emergent misalignment is influenced by training dynamics, model priors, and data.
problem Emergent misalignment in models
method Exploring training dynamics, model priors, and data
result Activation deltas before and after narrow fine-tuning correlate with their similarities when measured with the last prompt-token activations.
Paper calculates topological complexity of robot movement in narrow aisles.
problem Determining minimum number of scenarios for robot movement in a narrow strip.
method Examined cohomology ring of ordered configuration space to find lower bound.
result Lower bound for minimum number of cases in robot movement program.
Neural networks can approximate functions uniformly across various measures.
problem Universal approximation of functions across different probability measures.
method Proving neural networks are dense in Orlicz spaces, extending classical theorems.
result Neural networks uniformly approximate functions for weakly compact families of measures.
We introduce a deep multitask architecture to integrate multityped representations of multimodal objects. This multitype exposition is less abstract than the multimodal characterization, but more machine-friendly, and thus is more precise to model. For example, an image can be described by multiple visual views, which …
Tilting loss functions improves machine learning performance.
problem Improving machine learning models, especially in under- and over-parameterized networks.
method Using evolving loss functions that emphasize different classes cyclically.
result Dynamical loss functions lead to better generalization and stability in training.
Study geodesic orbits on noncompact curved spaces, proving their distribution and counting.
problem Counting and equidistribution of periodic orbits on noncompact manifolds.
method Proved equidistribution in narrow topology, deduced exact asymptotic counting.
result Exact asymptotic counting of periodic orbits on noncompact manifolds.
In (\cite{zhang2014nonlinear,zhang2014nonlinear2}), we have viewed machine learning as a coding and dimensionality reduction problem, and further proposed a simple unsupervised dimensionality reduction method, entitled deep distributed random samplings (DDRS). In this paper, we further extend it to supervised learning …
Adversarial attacks reduce deep learning beam selection performance in mmWave 5G.
problem Adversarial attacks degrade deep learning-based beam selection in mmWave 5G.
method Generates adversarial perturbations to RSS inputs to manipulate DNN predictions.
result Significant reduction in IA performance due to adversarial perturbations.
Machine learning models predict crash rates on narrow lanes.
problem Impact of narrow lanes on arterial road vehicle crashes.
method Applied random forest and least squares boosting machine learning algorithms to crash data.
result Random forest model identified as best for studying narrow lanes' safety impact.
This paper considers the generation of prediction intervals (PIs) by neural networks for quantifying uncertainty in regression tasks. It is axiomatic that high-quality PIs should be as narrow as possible, whilst capturing a specified portion of data. We derive a loss function directly from this axiom that requires no d…
HLOB predicts mid-price changes in L.O.Bs using deep learning.
problem Forecasting mid-price changes in Limit Order Books.
method HLOB uses a deep learning model with an Information Filtering Network and Homological Convolutional Neural Networks.
result HLOB outperforms state-of-the-art models in real-world datasets.
Compression affects deep networks differently, impacting underrepresented data points.
problem Disparate impact of compression on different classes and images.
method Analysis of deep neural network pruning and quantization effects.
result Compression disproportionately impacts model performance on underrepresented data points.
Wide networks avoid bad local minima, proving phase transition.
problem Understanding the optimization landscape of neural networks.
method Proving phase transitions in loss surfaces of wide and narrow networks.
result Phase transition from narrow to wide networks: no bad local minima.
Study shows DNNs can recover functions with fewer samples than model parameters at overparameterization.
problem Determining reliable function recovery in overparameterized deep neural networks.
method Introducing 'local linear recovery' (LLR) and proving upper bounds on sample sizes for recovery.
result Upper bounds on optimistic sample sizes for function recovery in overparameterized DNNs are achieved.
LAMVI-2 visualizes word embedding model tuning for developers.
problem Tuning deep learning models is complex and time-consuming.
method Introduces LAMVI-2, a visual analytics system for comparing hyperparameter settings.
result LAMVI-2 helps developers quickly and accurately choose effective models.
Unified approach to continual learning using generative replay and open set recognition.
problem Catastrophic interference and recognition of out-of-distribution data in deep neural networks.
method Probabilistic approach based on variational inference in a deep autoencoder model, using generative replay and open set recognition.
result The approach significantly alleviates catastrophic interference and distinguishes out-of-distribution data.
New algorithm ensures global convergence in deep neural networks beyond NTK regime.
problem Existing global convergence guarantees do not apply to practical deep networks.
method Proposes an algorithm with global convergence guarantees under the expressivity condition.
result Algorithm ensures global convergence in practical settings beyond NTK regime.
Proposes efficient training method for deep thin networks.
problem Deploying deep learning models with accuracy and compactness.
method Three-stage method: widen, warm up, fine tune.
result Deep thin networks trained with method outperform standard deep networks.
This paper presents a method to automatically generate high-quality prediction intervals for neural networks.
problem Accurate uncertainty quantification for deep learning models in real-world applications.
method Dual neural network approach with a novel loss function to balance prediction interval width and coverage.
result Our method produces significantly narrower prediction intervals with higher probability coverage compared to state-of-the-art methods.
We introduce a learning-based framework to optimize tensor programs for deep learning workloads. Efficient implementations of tensor operators, such as matrix multiplication and high dimensional convolution, are key enablers of effective deep learning systems. However, existing systems rely on manually optimized librar…
DOFEN improves DNN performance on tabular data benchmarks.
problem DOFEN tackles the performance gap between DNNs and tree-based models on tabular data.
method DOFEN uses a two-level rODT forest ensembling process inspired by oblivious decision trees.
result DOFEN achieves state-of-the-art results on the Tabular Benchmark.
New model explains deep learning performance at large learning rates.
problem Understanding deep learning performance at different learning rates.
method Developed neural networks with solvable training dynamics.
result Large learning rates lead to convergence to flatter minima.
BackPACK extends PyTorch to compute additional gradient info.
problem Lack of efficient tools for computing mini-batch variance and Hessian approximations.
method BackPACK builds on PyTorch to automatically compute additional derivatives.
result BackPACK enables efficient computation of various derivative quantities.
This paper compares deep transfer learning with classical ML in low-shot text classification.
problem Low-shot text classification with limited labeled data.
method Comparison of BERT and top classical ML approaches on a sentiment classification task.
result BERT outperforms classical ML by 9.7% on average with 100 labeled examples per class.
Empirical analysis of gradient descent optimizers in Deep RL.
problem Performance degradation in gradient descent methods for Deep RL.
method Analysis of various gradient descent optimizers and their hyperparameters.
result Adaptive optimizers have a narrow effective learning rate window, diverging in other cases.