PF-RNNs use particle filtering to model uncertainty in RNNs for better sequential data prediction.
problem Highly variable and noisy sequential data.
method PF-RNNs maintain a latent state distribution as a set of particles, updating with Bayes rule.
result PF-RNNs outperform standard RNNs on various sequence prediction tasks.
Compact RNNs reduce parameters and improve efficiency.
problem High computational cost of RNNs with large inputs.
method Block-Term Tensor Decomposition (BT-TD) to reduce RNN parameters.
result BT-RNN achieves better accuracy and faster convergence than standard RNNs.
R2N2 combines linear VAR and nonlinear RNN models for multivariate time series.
problem Multivariate time series modeling with poor predictive performance or complex models.
method R2N2 (Residual RNN) combines simple linear VAR and complex RNN models.
result R2N2 outperforms VAR and RNN alone, and is faster to train.
Diagonal RNNs improve music modeling performance and speed.
problem Improving symbolic music modeling efficiency and accuracy.
method Introduced diagonal recurrent matrices in RNNs for music modeling.
result Diagonal RNNs achieve better test likelihood and faster convergence.
ARNN augments RNNs with user-contextual preference for better session-based recommendations.
problem Limited context-awareness in RNN session models.
method Proposes ARNN that uses PNN to extract high-order user-contextual preference.
result ARNN outperforms baseline RNN by a large margin with rich user-side contexts.
Interpretable RNN uses sparse recovery for better performance.
problem Interpreting the internal workings of RNNs.
method Sequential Sparse Recovery + SISTA algorithm.
result SISTA-RNN achieves better performance and is more interpretable.
Three GRU variants reduce parameters in RNNs, improving efficiency.
problem Reducing computational expense in RNNs.
method Three variants of GRU with reduced parameters in update and reset gates.
result Variant models perform similarly to original GRU RNN models.
New model handles uneven time intervals better than traditional methods.
problem Irregularly-sampled time series data.
method Generalizes RNNs to ODE-RNNs, explicitly modeling observation gaps.
result ODE-RNNs outperform traditional models on irregular data.
Combines gating and tensor products for RNNs to improve performance.
problem Improving RNNs' ability to capture long-term dependencies.
method Proposes a novel RNN architecture combining gating mechanism and tensor products.
result Significant performance improvement on word-level and character-level language modeling tasks.
In this paper, we explore different ways to extend a recurrent neural network (RNN) to a \textit{deep} RNN. We start by arguing that the concept of depth in an RNN is not as clear as it is in feedforward neural networks. By carefully analyzing and understanding the architecture of an RNN, however, we find three points …
Paper compresses RNNs using HT decomposition for better performance.
problem Large model sizes of RNNs in sequence analysis.
method Hierarchical Tucker (HT) tensor decomposition for model compression.
result HT-LSTM achieves better compression and accuracy than state-of-the-art methods.
This paper proposes structurally sparse RNNs to reduce computational and memory costs.
problem Heavy computational and memory burden in fully connected RNNs.
method Study structurally sparse RNNs, reducing recurrent operations and weights.
result Structurally sparse RNNs achieve competitive performance with reduced costs.
Novel method improves training RNNs by accelerating gradient descent.
problem Vanishing and exploding gradient problems in RNNs training.
method Adaptive stochastic Nesterov accelerated quasi-Newton method.
result Improved performance in training RNNs with low per-iteration cost.
Transfer learning improves clinical time series prediction with limited data.
problem Training deep RNNs for clinical tasks requires large labeled data and tuning.
method Transfer learning from pre-trained RNNs on multiple tasks to new tasks.
result Features from pre-trained RNNs improve model performance and robustness.
Noisin injects random noise into RNNs to prevent overfitting.
problem Overfitting in RNNs for sequential data.
method Injects random noise into RNN hidden states and maximizes marginal likelihood.
result Noisin improves language model performance by up to 12.2% on Penn Treebank.
GORU combines unitary and gated RNNs for better long-term memory management.
problem Learning to effectively manage long-term memory in neural networks.
method Extending unitary RNNs with a gating mechanism to forget irrelevant information.
result GORU outperforms LSTMs, GRUs, and Unitary RNNs on long-term dependency tasks.
In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pai…
Recurrent Neural Networks (RNNs) have long been recognized for their potential to model complex time series. However, it remains to be determined what optimization techniques and recurrent architectures can be used to best realize this potential. The experiments presented take a deep look into Hessian free optimization…
Paper proposes quantizing RNNs to save space and power.
problem Over-parameterization in RNNs leads to inefficiency.
method Increases bit-width reduction for accuracy preservation.
result RNNs can maintain accuracy with reduced precision.
A new model combines spatial and spectral features for HSI classification.
problem Inefficiency and difficulty in training RNNs for HSI classification.
method Proposes St-SS-pGRU combining shorten RNN, converlusion layer, and parallel-GRU.
result Better performance and robustness in HSI classification.
Proposes Fusion Recurrent Neural Network for sequence data.
problem Improving sequence learning for practical applications.
method Fusion module and Transport module for sequence data.
result Fusion RNN performs comparably to state-of-the-art RNNs.
Delayed-RNN approximates stacked and bidirectional RNNs.
problem Improving RNN expressiveness and representational capacity.
method Weight-constrained delayed-RNN, equivalent to stacked-RNNs, with partial acausality.
result Delayed-RNN can approximate stacked and bidirectional RNNs, outperforming them in some tasks.
Study shows challenges in converting RNNs to FSMs due to computational complexity.
problem Understanding the equivalence and distance between RNNs and FSMs.
method Computational proofs for equivalence and distance problems between RNNs and FSMs.
result Undecidability and hardness of approximation problems between RNNs and FSMs.
FastGRNN improves RNN accuracy while drastically reducing model size.
problem Inaccurate training and inefficient prediction in RNNs.
method FastGRNN uses a residual connection and gate to achieve state-of-the-art accuracy with a much smaller model.
result FastGRNN achieves state-of-the-art accuracy with models up to 35x smaller than existing RNNs.
RNN(p) improves power consumption forecasts with interpretable models.
problem Improving power consumption forecasts for energy sector decisions.
method RNN(p) models with p time lags, using structured feedbacks.
result RNN(p) models achieve excellent forecasting accuracy and interpretability.
Anticipation-RNN generates music with user-defined constraints.
problem Interactive music generation with user-defined constraints.
method Introduces Anticipation-RNN, a novel architecture that allows user-defined positional constraints.
result Demonstrates efficient generation of melodies satisfying user-defined constraints.
Paper questions RNN and LSTM's long-term memory and introduces a new definition.
problem Whether RNN and LSTM have long-term memory.
method Introduced a new definition of long-term memory and modified RNN and LSTM to test it.
result RNN and LSTM do not meet the new definition of long-term memory.
GPU-optimized ES-RNN boosts time series forecasting speed by 322x.
problem Efficiently forecasting time series data.
method Vectorized GPU implementation of ES-RNN.
result Up to 322x speedup in training time.
RNN-HAR model improves VaR forecasting with long-memory and non-linear dynamics.
problem Efficiently forecasting Value at Risk (VaR) with long-memory and non-linear realized volatility.
method Loss-based generalized Bayesian inference with Sequential Monte Carlo for model estimation and prediction.
result RNN-HAR model consistently outperforms other VaR forecasting models.
Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs. Unlike feedforward neural networks, RNNs have cyclic connections making them powerful for modeling sequences. They have been successfully u…
Flexible DSL generates novel RNN architectures for various tasks.
problem Limited flexibility and components in existing RNN architectures.
method Domain-specific language (DSL) for automated architecture search.
result Novel RNN architectures perform well on language modeling and machine translation tasks.
Bayesian RNN model forecasts and quantifies uncertainty in spatio-temporal data.
problem Uncertainty quantification in nonlinear spatio-temporal systems.
method Developed a Bayesian RNN model to forecast and quantify uncertainty rigorously.
result The model maintains forecast accuracy while quantifying uncertainty formally.
Mixture Density RNNs learn to model predictions as multiple Gaussian distributions.
problem Understanding predictions made by MD-RNNs in complex environments.
method Analyzed predictions from trained MD-RNNs, focusing on their Gaussian components.
result MD-RNNs' Gaussian components separately model stochastic events and scenarios governed by different rules.
Frequentist method estimates uncertainty in RNNs without altering architecture.
problem Uncertainty quantification in RNNs for decision-making.
method Jackknife resampling and influence functions to estimate variability.
result The method provides theoretical coverage guarantees on uncertainty intervals.
Deep Neural Network (DNN) acoustic models have yielded many state-of-the-art results in Automatic Speech Recognition (ASR) tasks. More recently, Recurrent Neural Network (RNN) models have been shown to outperform DNNs counterparts. However, state-of-the-art DNN and RNN models tend to be impractical to deploy on embedde…
This study improves knowledge distillation for RNN-T models with noisy labels.
problem Challenges in distilling knowledge from RNN-T models with variable quality teachers.
method Full-sum distillation and sequence-level knowledge distillation.
result Full-sum distillation outperforms other methods for RNN-T models, especially for bad teachers.
A new RNN model tackles long-time dependencies with fast, invertible, and memory-efficient hidden states.
problem Challenges in processing sequential inputs with long-time dependencies in RNNs.
method A novel RNN architecture based on a Hamiltonian system of oscillators.
result The proposed RNN mitigates exploding and vanishing gradient problems, providing state-of-the-art performance.
Predicts individual septic shock children's vasoactive response using RNN.
problem Personalized physiologic responses to vasoactive titrations in septic shock children.
method Retrospective analysis of EMR data using a Recurrent Neural Network (RNN).
result RNN model predicted physiologic responses more accurately than a linear model.
In this paper, we propose a novel neural network model called RNN Encoder-Decoder that consists of two recurrent neural networks (RNN). One RNN encodes a sequence of symbols into a fixed-length vector representation, and the other decodes the representation into another sequence of symbols. The encoder and decoder of t…
ASARS integrates temporal dynamics from CF into session-based RNN for better personalized recommendations.
problem Discarding long-term data across sessions in session-based RNNs.
method ASARS framework that combines attentional network and inter-session temporal dynamic model.
result ASARS improves personalized recommendation performance on four real datasets.
AC-RNN improves RNN for sequence labeling tasks.
problem RNN's exposure bias in maximum-likelihood training.
method Actor-Critic training for RNNs.
result AC-RNN outperforms CRF on NER and CCG tagging.
Develops a tensor network framework to reduce RNN complexity for high-dimensional sequence modeling.
problem Exponential parameter growth in RNNs for large multidimensional data.
method Embeds a multi-linear graph filter in a tensor network architecture to approximate RNN hidden states.
result Demonstrates superior performance and reduced complexity compared to traditional RNNs.
Finite precision RNNs have varying computational power, with LSTMs and ReLU-RNNs being more powerful.
problem Understanding the computational limits of finite precision RNNs for language recognition.
method Comparison of different RNN variants with finite precision and linear computation time.
result LSTMs and ReLU-RNNs are strictly stronger than other RNN variants in terms of computational power.
PC-RNN reconstructs MRI images from undersampled data with more details.
problem Recovering fine details from undersampled MRI data.
method Pyramid Convolutional RNN (PC-RNN) with three ConvRNN modules for multi-scale reconstruction.
result PC-RNN outperforms other methods in recovering more details from MRI images.
New modifiers improve noisy RNN replay in hippocampal networks.
problem Improving noisy RNN replay in hippocampal networks.
method Three approaches: hidden state leakage, adaptation, and momentum.
result Hidden state leakage, adaptation, and momentum improve noisy RNN replay.
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
The paper develops a convex parameterization for robust RNNs ensuring stability and robustness.
problem Lack of stability and robustness guarantees in RNNs for sequence-to-sequence mapping applications.
method Formulated convex sets of RNNs with stability and robustness guarantees using incremental quadratic constraints.
result The proposed model structure ensures global exponential stability and bounds on incremental ℓ2 gain. AntMan compresses RNNs for faster inference with minimal accuracy loss.
problem Inference performance, cost, and memory requirements of complex RNN models.
method Structured sparsity combined with low-rank decomposition.
result Up to 100x computation reduction with less than 1pt accuracy drop.