Paper proposes dual recurrent attention units for VQA models.
problem Comprehending visual and textual data for accurate question answering.
method Introduces and evaluates recurrent attention mechanisms in VQA models.
result Dual Recurrent Attention Units (RAUs) improve VQA performance.
AReLU uses attention-based rectification to improve neural network performance.
problem Improving neural network performance through better activation functions.
method Integrates attention mechanism with rectified linear unit (ReLU) to learn and scale feature maps.
result AReLU significantly boosts performance of most network architectures with minimal changes.
RAU integrates attention into GRU for better sequence learning.
problem Lack of attention mechanism in GRU leads to information redundancy or loss.
method RAU adds an attention gate to GRU to adaptively focus on regions of interest.
result RAU consistently outperforms GRU and other methods in various tasks.
Improved speech recognition with end-to-end attention models.
problem End-to-end speech recognition with open-vocabulary.
method Sequence-to-sequence attention-based models on subword units, new pretraining scheme, CTC loss function, LSTM language models.
result State-of-the-art word error rates (3.54% and 3.82%) on LibriSpeech test-clean.
A new method for optimizing regression problems with ReLU units converges.
problem Optimizing regression problems involving ReLU units in large language models.
method Introduced a greedy algorithm based on approximate Newton method, proving convergence in terms of the distance to optimal solution.
result The method converges in the sense of the distance to optimal solution under certain assumptions.
Deep neural network learns hierarchical language family structure.
problem Language family dependency and external information affect classification accuracy.
method Hierarchical attentive units learn auxiliary tasks for robust internal representation.
result Improved classification accuracy on small and big language corpora.
Study examines dependence properties of Bayesian neural network units in finite-width networks.
problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.
RALM extends exposure for long-tail contents in real-time recommender systems.
problem Lack of timely exposure for long-tail contents in recommender systems.
method Real-time attention-based look-alike model (RALM) with seeds-to-user similarity prediction and user representation learning.
result RALM outperforms traditional look-alike models in effectiveness and real-time performance.
A deep learning model improves pedestrian tracking accuracy.
problem Pedestrian tracking accuracy is low, especially with inertial measurement unit.
method Deep learning model using IMU and LIDAR data, attention mechanism.
result Preliminary results show improved accuracy.
ATFM predicts traffic flow using attention mechanism and neural networks.
problem Predicting traffic flow with diverse factors integration.
method Unified neural network (ATFM) with attention mechanism and ConvLSTM units.
result ATFM outperforms in predicting citywide short-term/long-term traffic flow.
Study on hidden units in finite Bayesian neural networks and their tail properties.
problem Understanding the behavior of hidden units in finite Bayesian neural networks.
method Introduced a generalized Weibull-tail property to describe hidden units tails.
result Unit priors become heavier-tailed going deeper, providing insights into finite Bayesian neural networks.
Study predicts blood pressure response to fluid bolus therapy with high accuracy.
problem Predicting successful response to fluid bolus therapy in hypotensive ICU patients.
method Used attention-based LSTM and GRU neural networks on a large ICU database.
result Stacked LSTM with attention mechanism achieved highest accuracy of 0.852.
Neural network predicts cardiovascular events from EHRs with high accuracy.
problem Predicting onset of cardiovascular diseases from electronic health records.
method Multi-task gated recurrent units with attention mechanism.
result Model outperforms clinical risk scores in predicting stroke and myocardial infarction.
Paper presents a neural network method for automated bug and ticket classification.
problem Automated classification of bug and ticket content in systems.
method Recurrent neural network with hierarchical attention mechanism.
result The method outperforms previous approaches on two datasets.
A new KD method distills dataset-based knowledge using MHA.
problem Distilling knowledge from large teacher networks to small student networks.
method Graph-based knowledge distillation by multi-head attention network.
result The method improves SN performance by 7.05% on CIFAR100.
Modeling student behaviors and multiple predictions for early intervention.
problem Predicting student outcomes and interactions among multiple tasks.
method Proposes a variant of LSTM and soft-attention mechanism for heterogeneous behaviors, and co-attention mechanism for task interactions.
result Demonstrated effectiveness in predicting student outcomes and interactions.
RUM improves RNN's long-term memory by using unitary matrices.
problem Limited capacity of RNN to manipulate long-term memory.
method Proposes Rotational Unit of Memory (RUM) with unitary matrices.
result RUM learns long-term dependencies and improves state-of-the-art results.
A3T-GCN improves traffic forecasting by capturing spatial and temporal dependencies.
problem Accurate real-time traffic forecasting in complex road networks.
method Attention Temporal Graph Convolutional Network (A3T-GCN) integrating recurrent units and graph convolutional network.
result Improved prediction accuracy through attention mechanism and global temporal information.
TSAM predicts directed temporal links using GCN and self-attention.
problem Predicting links in directed temporal networks.
method GCN, self-attention mechanism, autoencoder architecture, graph attentional layers, graph convolutional layers, graph recurrent unit layer.
result TSAM outperforms benchmarks on four realistic networks.
New approach reduces model size for Transformer architectures.
problem Large embedding dimensions limit model applicability.
method Identified low-rank bottleneck in multi-head attention.
result Reducing head size to sequence length improves model performance.
Unified perspective on Hopfield networks with attention module.
problem Understanding and optimizing Hopfield networks with attention mechanisms.
method Study of BM counterparts of modern Hopfield networks and their salient properties.
result Introduction of AttnBM with tractable likelihood and gradient.
Study compares deep learning models for traffic forecasting, highlighting graph elements' impact.
problem Challenges in forecasting spatial-temporal traffic patterns.
method In-depth comparative study of four deep neural network models with different basic elements.
result Graph attention improves long-term predictions in traffic forecasting models.
This paper improves multichannel speech enhancement using complex ratio masking and channel-attention.
problem Limited performance of deep learning methods in multichannel speech enhancement.
method Introduces complex ratio masking and channel-attention mechanism inside a U-Net architecture.
result Demonstrates superior performance on the CHiME-3 dataset.
Deep learning predicts vascular disease from medical history.
problem Predicting high-risk vascular diseases from medical records.
method Medical History-based Prediction using Attention Network (MeHPAN) models.
result MeHPAN models outperform standard classification models.
Deep learning models forecast stock market orders over multiple time frames.
problem Forecasting stock market orders over varying time frames.
method Encoder-decoder models with sequence-to-sequence and Attention mechanisms, leveraging Intelligent Processing Units (IPUs) for faster training.
result Multi-horizon forecasting outperforms single-horizon models, especially for long prediction periods.
Improved stock price prediction using LSTM with customized loss function and analyst calls.
problem Forecasting stock prices using LSTM neural networks.
method Customized LSTM model with improved loss function, analyst calls integration, and attention units.
result Improved performance of LSTM model over ARIMA for stock price prediction.
Attention models outperform RNNs in clinical time series analysis.
problem Clinical time series analysis using deep learning models.
method Developed SAnD architecture using masked self-attention mechanism, positional encoding, and dense interpolation.
result Achieved state-of-the-art performance in all tasks on MIMIC-III benchmark datasets.
Study improves stock price prediction using advanced ML models.
problem Improving financial forecasting accuracy in stock markets.
method Evaluation of RNN architectures including LSTM, GRU, and attention-based models.
result Attention-based models outperform others in capturing complex dependencies.
HAGAN uses hierarchical attention to improve cross-domain sentiment classification.
problem Cross-domain sentiment classification with domain discrepancy.
method Hierarchical attention in GANs to produce domain-indistinguishable document representations.
result HAGAN outperforms existing methods on Amazon review dataset.
Convolutional-deconvolution networks can be adopted to perform end-to-end saliency detection. But, they do not work well with objects of multiple scales. To overcome such a limitation, in this work, we propose a recurrent attentional convolutional-deconvolution network (RACDNN). Using spatial transformer and recurrent …
Paper proposes a graph neural network for accurate long-term ILI prediction.
problem Limited long-term prediction performance and spatio-temporal dependency in existing models.
method Cross-location attention based graph neural network (Cola-GNN) for time series embeddings and location aware attentions.
result Proposed method shows strong predictive performance and interpretable results for long-term epidemic predictions.
GOBO compresses 99.9% of BERT model parameters to 3 bits, improving inference efficiency.
problem Efficient execution of attention-based NLP models, especially in terms of latency and energy consumption.
method GOBO quantizes 32-bit floating-point parameters to 3 bits without fine-tuning, using hardware compression and co-designed architectures.
result GOBO maintains model accuracy while significantly reducing inference latency and energy consumption.
PASS model predicts disease progression with both accuracy and interpretability.
problem Balancing accurate disease prediction with clinically interpretable models.
method Phased LSTM units with attention mechanism for non-stationary state dynamics.
result PASS model achieves superior predictive accuracy and interpretable representations.
CLVSA predicts financial market trends using LSTM and attention mechanisms.
problem Predicting trends in financial markets due to complex interactions.
method Hybrid model combining LSTM, sequence-to-sequence, attention, and convolutional LSTM.
result CLVSA outperforms basic models in predicting financial market trends.
Efficient neural network optimization reduces costs and improves model performance.
problem High computational costs in optimizing neural networks, especially at scale.
method Introduces self-attentive feed-forward neural units (SAFFU) for efficient optimization.
result Explicit solutions outperform models optimized by backpropagation alone, and further training with backpropagation leads to better optima from smaller data sets.
One important effect of price shocks in the United States has been increased political attention paid to the structure and performance of oil and natural gas markets, along with some governmental support for energy conservation. This paper describes how price changes helped lead the emergence of a political agenda acco…
Amobee won 3rd and 1st place in SemEval 2018 sentiment classification tasks.
problem Sentiment classification in multiple languages.
method Training GRU-CNN model with word embeddings and stacking ensembles.
result 3rd and 1st place in valence ordinal classification sub-tasks in English and Spanish.
AuGMEnT network struggles with long-term memory for hierarchical tasks.
problem Learning and memory in neural networks, especially hierarchical tasks.
method Introduced hybrid AuGMEnT with leaky and non-leaky memory units.
result Hybrid AuGMEnT solves hierarchical and distractor tasks.
ACFM predicts crowd flow adaptively integrating various factors.
problem Adaptive integration of factors affecting crowd flow changes.
method Unified neural network module with attention mechanism.
result Significant improvements over state-of-the-art methods.
The paper uses GRU and self-attention for SPY option pricing.
problem Precise prediction of SPY option prices for better investment decisions.
method Partitioned dataset, built four models, used SHAP for interpretation.
result Self-attention GRU model outperforms traditional models.
Deep learning models predict ICU readmission with varying accuracy.
problem Predicting ICU readmission risk using deep learning architectures.
method Several deep learning architectures including attention-based models, recurrent layers, neural ODEs, and embeddings were trained on MIMIC-III data.
result Attention-based models with neural ODEs achieved highest predictive accuracy.
MPSA-DenseNet improves accent classification accuracy.
problem Accurate English accent identification.
method Combines multi-task learning and PSA attention mechanism with DenseNet.
result MPSA-DenseNet outperforms other models in accent classification.
Paper shows how LSTM can remember long sequences by attending to persisted information.
problem LSTMs struggle with long sequences due to fading information and bias towards recent data.
method The paper introduces a mechanism that allows LSTMs to attend to information in memory based on how long it was persisted by the gating mechanism.
result The method improves LSTM's ability to process long sequences by retrieving information proportionally to its persistence in memory.
Study on deep multi-head self-attention dynamics, proving homogenized limits under specific scalings.
problem Understanding the behavior of deep multi-head self-attention models as depth increases.
method Random model of deep multi-head self-attention, viewing depth as time, and analyzing the residual stream as a particle system.
result Homogenized limit of the dynamics, leading to deterministic or stochastic behavior depending on scaling, with implications for representation collapse.
Binary testing for softmax models requires many samples, similar to leverage score models.
problem Binary hypothesis testing for softmax models and leverage score models.
method Analyzing sample complexity and drawing analogies between models.
result Sample complexity is asymptotically \(O(ε^{-2})\), where \(ε\) is the distance between model parameters.
Integrates global information into dropout for better text classification.
problem Improving neural networks for text classification.
method GI-Dropout, a novel dropout method integrating global information.
result Demonstrates the effectiveness of GI-Dropout on seven text classification tasks.
MONet learns to decompose scenes into meaningful components without supervision.
problem Learning meaningful scene decompositions without labeled data.
method MONet combines a VAE and recurrent attention network to learn decompositions of 3D scenes.
result MONet can learn to represent 3D scenes into meaningful components like objects and background.
Automates detecting problem statements in peer assessments.
problem Identifying problem statements in peer assessment reviews.
method Used machine learning models including neural networks and traditional classifiers.
result Hierarchical Attention Network classifier achieved 93.1% accuracy.