ForecastNet uses a time-variant deep feed-forward neural network for better multi-step-ahead time series forecasting.
problem Time-invariant architectures limit multi-step-ahead forecasting.
method ForecastNet employs a deep feed-forward architecture with time-variant parameters and interleaved outputs.
result ForecastNet outperforms other models on multi-step-ahead time series forecasting tasks.
Proposes RMN for learning long-term dependencies in feed-forward networks.
problem Complicated training of deep RNN architectures.
method Residual Memory Neural Network (RMN) with residual and time-delayed connections.
result RMN and BRMN outperform LSTM and BLSTM networks in learning long-term and hierarchical information.
This work optimizes reservoir computing models by linking recurrence and non-linear dynamics.
problem Understanding how recurrence and non-linear dynamics in cortical networks contribute to their function.
method Transformed time-continuous, recurrent dynamics into an effective feed-forward structure of linear and non-linear temporal kernels.
result Optimal time-series classifiers can be built from random reservoir networks, demonstrating significant performance gains.
New architecture avoids forgetting old tasks while learning new ones.
problem Catastrophic forgetting in neural networks when learning multiple tasks.
method Sequential learning paradigm with reusing features from previous tasks.
result Reduces forgetting and requires fewer resources for new task learning.
Deep network learns Obstacle Tower challenge without human demonstrations.
problem Master procedurally generated levels that get progressively harder.
method Deep Reinforcement Learning with a simple feed-forward network.
result Performed competitively in a reinforcement learning competition.
The study connects deep neural networks with statistical mechanics, revealing natural activation functions.
problem Understanding the activation functions in deep neural networks.
method Statistical Mechanics model of deep neural networks, focusing on encoding, validation, and propagation steps.
result A set of natural activations including Sigmoid, tanh, ReLU, and Swish are identified.
New study shows neural networks need many samples for training.
problem How much data is needed to train a ReLU feed-forward neural network?
method Theoretical and empirical analysis of ReLU feed-forward neural networks.
result Generalization error scales at 1 / n 1/\sqrt{n} 1/ n in sample size n n n . Proposes a new CG interpretation of neural networks for better theoretical analysis.
problem Lack of theoretical analysis in neural networks interpretation.
method Interprets neural networks as chain graphs and feed-forward as approximate inference.
result Provides novel theoretical support and insights for various neural network techniques.
Stable processes emerge as limits of deep neural networks with symmetric stable distributions.
problem Understanding the behavior of deep neural networks as they become infinitely wide.
method Analyzing fully connected feed-forward deep neural networks with symmetric stable distributions and showing the limit as a stable process.
result The infinite wide limit of the network is a stable process with multivariate stable distributions.
Neural nets trained with linear discriminant initialization converge faster and more accurately.
problem Training feed-forward neural networks efficiently and accurately.
method Initialize first layer weights with linear discriminants.
result Asymptotic higher accuracy and faster convergence.
We introduce Dynamic Deep Neural Networks (D2NN), a new type of feed-forward deep neural network that allows selective execution. Given an input, only a subset of D2NN neurons are executed, and the particular subset is determined by the D2NN itself. By pruning unnecessary computation depending on input, D2NNs provide a…
Proposes a feed-forward method for uncertainty propagation in neural networks.
problem Uncertainty in neural network inputs and activations.
method Feed-forward inference method for belief and neural networks, considering both mean and variance.
result Propagates uncertainty through neural network layers and parameters.
We present experiments demonstrating that some other form of capacity control, different from network size, plays a central role in learning multilayer feed-forward networks. We argue, partially through analogy to matrix factorization, that this is an inductive bias that can help shed light on deep learning.
DeepPhaseCut uses neural networks to improve Fourier phase retrieval.
problem Fourier phase retrieval from magnitude data.
method Unsupervised feed-forward neural network with cycleGAN training.
result Outperforms existing methods in Fourier phase retrieval.
NN2Poly converts deep neural networks into polynomial models for better understanding.
problem Improving neural network interpretability and theoretical understanding.
method Taylor expansion on activation functions, combinatorial properties, and polynomial coefficients calculation.
result NN2Poly accurately represents deep feed-forward neural networks as polynomial models.
New bound for neural nets on non-iid data.
problem Generalization of deep nets for dependent data.
method Establishes a generalization bound for feed-forward neural networks on φ φ φ -mixing data. result Proves neural nets can generalize well on non-iid data.
This paper removes the finite variance assumption for deep convolutional neural networks.
problem Removing the finite variance assumption for deep convolutional neural networks.
method Assuming iid parameters distributed according to a stable distribution, the paper shows that the infinite-channel limit of a deep feed-forward convolutional neural network is a multivariate stable stochastic process.
result The infinite-channel limit of a deep feed-forward convolutional neural network, under suitable scaling, is a multivariate stable stochastic process.
Feed-forward nets fail to learn equality relations, but adding DR units helps.
problem Feed-forward neural networks struggle to learn equality relations reliably.
method Introduced differential rectifier (DR) units to create an inductive bias.
result DR units enable feed-forward nets to learn equality relations reliably.
Single-pass method estimates neural network uncertainty.
problem Uncertainty estimation in deep learning requires multiple passes.
method Probabilistic reasoning over neural network depths.
result Single forward pass for uncertainty estimation.
A new algorithm speeds up neural network derivative calculations.
problem Exponential runtime of autodifferentiation for high-order derivatives in neural networks.
method n-TangentProp, a quasilinear algorithm for computing higher-order derivatives.
result Computes exact derivatives in quasilinear time, not exponential.
We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requirements. The max-pooling loss training can be further guided by initializing with a cross-entropy loss trained network. A posterior smoothin…
Adaptive Group Lasso selects important features in neural networks.
problem Lack of interpretability in neural networks.
method Adaptive Group Lasso for feature selection.
result Consistent feature selection for neural networks with theoretical guarantee.
Stable recurrent models perform similarly to unstable ones, proving useful for sequence tasks.
problem The lack of stability in recurrent neural networks hinders their practical application.
method Theoretical analysis and empirical testing of stable recurrent models.
result Stable recurrent models can be well approximated by feed-forward networks, performing similarly to unstable ones.
Paper characterizes and constructs universal approximators for neural networks.
problem Limited understanding of universal approximation in neural networks.
method Characterization, representation, construction method, existence result for any universal approximator.
result Improved capabilities of feed-forward architecture to approximate continuous functions.
This paper improves speech synthesis using a DGP with SRU for naturalness.
problem Improving naturalness in synthetic speech.
method Deep Gaussian process with a recurrent architecture using SRU.
result SRU-DGP outperforms other models in naturalness of synthetic speech.
New deep architecture improves stability and robustness of neural networks.
problem Improving stability and robustness of neural networks.
method Proposes a new deep architecture using implicit residual blocks with fixed point definitions.
result Improved stability and generalization power with fewer hyperparameters.
Adapting robust statistics to neural networks, researchers found neural networks can be more robust with certain loss functions.
problem The robustness of neural networks in complex learning tasks.
method Adapting the regression breakdown point from robust statistics to neural networks and comparing different configurations and contamination settings.
result Neural networks can benefit from robust loss functions, as demonstrated in extensive simulations.
Efficient neural network optimization reduces costs and improves model performance.
problem High computational costs in optimizing neural networks, especially at scale.
method Introduces self-attentive feed-forward neural units (SAFFU) for efficient optimization.
result Explicit solutions outperform models optimized by backpropagation alone, and further training with backpropagation leads to better optima from smaller data sets.
SSFN self-estimates network size with low complexity and consistent performance.
problem Designing a self-estimating feed-forward network with low complexity and consistent performance.
method Joint optimization for layer and node estimation, low computational complexity, and use of lossless flow property and convex optimization.
result Consistent performance across Monte-Carlo trials and monotonically non-increasing cost with network growth.
New method uses neural networks to forecast spatial-temporal data.
problem Probabilistic forecasting of spatio-temporal data with causal structure.
method MMAF-guided learning with ensemble of stochastic feed-forward neural networks.
result Forecasting remains calibrated across multiple time horizons.
daForest improves deep forest ensemble for classification with better performance.
problem Improving classification accuracy with small training sets.
method SAMME.R boosting, feed-forward connection, hyper-parameters optimization.
result daForest outperforms neural networks and achieves state-of-the-art results.
New model adds persistent memory to self-attention layers for improved performance.
problem Improving transformer performance by removing feed-forward layers.
method Augmenting self-attention layers with persistent memory vectors.
result The model outperforms standard transformers on language modeling benchmarks.
Paper uses neural networks to analyze oil price impact on Iranian stock and industry indices.
problem Impact of oil price volatility on Tehran stock and industry indices.
method Feed-forward neural networks analysis of two periods: sanctions and post-sanctions.
result Neural networks predict stock and industry indices well, showing significant oil price volatility impact.
TFiLM expands convolutional models' receptive field with minimal overhead.
problem Capturing long-range dependencies in sequential data.
method A novel architectural component using a recurrent neural network to modulate convolutional model activations.
result TFiLM significantly improves learning speed and accuracy on various tasks.
The paper sets information-theoretic lower bounds for neural networks' parameter recovery and excess risk.
problem Establishing sample complexity lower bounds for neural network parameters and excess risk.
method Using information-theoretic tools, the paper proves lower bounds by constructing a generative network.
result Proves information-theoretic lower bounds for exact parameter recovery and positive excess risk.
New connection between FFNs and BNs improves performance and generalization.
problem Improving performance and generalization in classification and segmentation tasks.
method Characterizing FFNs as approximations of BNs and developing a new learning algorithm.
result Statistically learned BNs outperform FFNs in classification and segmentation tasks.
The Resilient Propagation (Rprop) algorithm has been very popular for backpropagation training of multilayer feed-forward neural networks in various applications. The standard Rprop however encounters difficulties in the context of deep neural networks as typically happens with gradient-based learning algorithms. In th…
Deep neural networks converge to Gaussian mixtures as layer width increases.
problem Understanding the distribution of outputs from deep neural networks.
method Proof and experiments with a simple model showing the convergence of neural network outputs to Gaussian mixtures.
result Neural networks converge to Gaussian mixtures as the width of the last hidden layer increases.
New artificial neuron doubles weight for improved deep learning accuracy.
problem Improving deep learning models' accuracy.
method Introducing a double-weight neuron, tested on MNIST and CIFAR-10 datasets.
result Significant improvement in classification accuracy for MNIST and CIFAR-10 datasets.
We investigate the capacity, convexity and characterization of a general family of norm-constrained feed-forward networks.
New method uses MMAF-guided learning for spatio-temporal probabilistic forecasts.
problem Probabilistic forecasting of spatio-temporal data with causal structure.
method Generalized Bayesian methodology, MMAF-guided learning, ensemble of stochastic feed-forward neural networks.
result Forecast performance comparable to, and sometimes better than, deep learning architectures.
Single-layer Transformer can approximate any sequence mapping.
problem Lack of theoretical understanding of Transformers.
method Review of linear algebra, probability, and optimization concepts; detailed analysis of Transformer architecture.
result A single-layer Transformer can approximate any continuous sequence-to-sequence mapping to arbitrary precision.
NFM improves deep learning by selectively processing hidden states.
problem Processing entire hidden states in each layer limits modularity and reusability.
method Introduces Neural Function Modules (NFM) with attention, sparsity, and feedback.
result Improves results in classification, generalization, generative modeling, and reinforcement learning.
New insights into how encoder-decoder networks generate attention matrices.
problem Understanding how encoder-decoder networks use attention matrices.
method Decomposing hidden states into temporal and input-driven components.
result Attention matrices are formed based on task requirements, not architecture type.
Deep LSTMs learn correlated representations of time-series data.
problem Learning nonlinear transformations and correlated embeddings of variable-length sequences.
method Use LSTMs to transform multi-view time-series data, then correlate outputs to find a fixed-dimensional representation.
result Deep LSTMs can effectively learn and project correlated representations of time-series data.
This paper optimizes autonomous vehicle controllers using data-driven methods.
problem Designing robust controllers for autonomous vehicles that handle external and internal disturbances.
method Data-driven approach using principal component analysis and time delay neural networks.
result Improved controller performance through a feed-forward compensator.
Mean field theory explains gradient backpropagation in deep dropout networks.
problem Understanding gradient backpropagation in deep dropout networks.
method Applied mean field theory to dropout networks, considering realistic training conditions.
result Gradient backpropagation length is limited by depth scales, not just independence assumption.
Lipschitz regularization improves neural network robustness by coupling weights across layers.
problem Improving neural network robustness under random input uncertainties.
method Regularization of neural networks by their Lipschitz constant, highlighting the coupling effect on weights across layers.
result Lipschitz regularization introduces a tradeoff between robustness and expressiveness, suggesting careful implementation.