Linear attention model outperforms full attention in language tasks.
problem Quadratic scaling of full attention limits model size.
method Introduced a linear attention mechanism.
result Linear attention models achieve comparable performance to full attention models.
Bayesian attention improves model performance and robustness.
problem Limited exploration of stochastic attention in neural networks.
method Introduces Bayesian attention belief networks using gamma and Weibull distributions.
result Outperforms deterministic and stochastic attention methods in accuracy and robustness.
GA-Net selectively attends to parts of a sequence for text classification.
problem Inefficient global attention mechanisms on long sequences.
method Gated Attention Network (GA-Net) using an auxiliary network to dynamically select and attend to important parts of the sequence.
result GA-Net achieves better performance with less computation and interpretability.
Paper proposes a method to interpret deep neural networks using attention mechanisms.
problem Interpreting deep neural network models to understand their performance.
method Proposes a novel method using attention mechanisms to analyze neural network models.
result Improved attention based method shows better classifier interpretation.
Enhances group convolutional networks with attention to learn meaningful relationships.
problem Lack of explicit means to learn meaningful relationships among symmetry patterns.
method Introduces attentive group equivariant convolutions, applying attention during convolution.
result Consistently outperforms conventional group convolutional networks on benchmark datasets.
Mathematical framework for understanding attention in neural networks.
problem Lack of theoretical understanding of attention in neural networks.
method Proposes a measure-theoretic model of attention and interprets self-attention as a system of self-interacting particles.
result Shows that attention is Lipschitz-continuous under suitable assumptions.
Two attention models improve human activity recognition by focusing on important signals and sensor modalities.
problem Noise and unimportant signal components in recurrent networks for human activity recognition.
method Temporal and sensor attention mechanisms with continuity constraints.
result State-of-the-art results on three datasets, showing improved understandability and mean F1 score.
A framework for transformer attention layers derived from SVR.
problem Developing principled attention mechanisms for transformers.
method Mapping self-attention to SVR, deriving new attention types.
result Improved transformer performance and efficiency.
Survey categorizes attention models across various domains.
problem Understanding and improving attention models in neural networks.
method Taxonomy and review of existing techniques.
result Provides a structured overview of attention models.
Self-attentive network improves emotion recognition in conversations.
problem Emotion recognition in dyadic conversations using deep learning.
method Introduces a novel self-attention mechanism for capturing temporal dynamics without a decoder.
result Outperforms state-of-the-art alternatives on the IEMOCAP benchmark.
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.
New model improves graph attention for relational data.
problem Improving graph attention models for relational data.
method Relational Graph Attention Networks (R-GAT) extending non-relational graph attention to relational data.
result R-GAT performs worse than expected, but some configurations marginally improve molecular property modeling.
Graph neural networks benefit from attention under specific conditions.
problem Understanding and improving the effectiveness of attention in graph neural networks.
method Designing controlled graph reasoning tasks, analyzing performance under various conditions, proposing weakly-supervised training.
result Attention can provide significant gains in performance under certain conditions, but its effect is often negligible or harmful.
AReLU uses attention-based rectification to improve neural network performance.
problem Improving neural network performance through better activation functions.
method Integrates attention mechanism with rectified linear unit (ReLU) to learn and scale feature maps.
result AReLU significantly boosts performance of most network architectures with minimal changes.
Analyzes self-attention in recurrent networks, proving it mitigates vanishing gradients.
problem Vanishing gradients in recurrent networks when capturing long-term dependencies.
method Formal analysis of self-attention's effect on gradient propagation, proposing a relevancy screening mechanism.
result Self-attention mitigates vanishing gradients in recurrent networks, providing guarantees.
Paper investigates Lipschitz constants of self-attention modules in neural networks.
problem Lipschitz constants of self-attention modules in neural networks.
method Proved standard dot-product self-attention is not Lipschitz for unbounded input domain. Proposed L2 self-attention that is Lipschitz. Derived upper bound on L2 self-attention's Lipschitz constant.
result Proved standard self-attention is not Lipschitz for unbounded input domain and proposed an alternative L2 self-attention that is Lipschitz.
Aligns attention distributions for improved accuracy and robustness.
problem Improving the accuracy and robustness of neural networks using attention mechanisms.
method Alignment attention that encourages key and query distributions to match within each head.
result Alignment attention leads to better accuracy, uncertainty estimation, and robustness across various tasks.
SpGAT learns graph representations using spectral attention for efficiency.
problem Efficiently capturing global graph patterns with minimal parameters.
method Introduces Spectral Graph Attention Network (SpGAT) using spectral domain attention mechanisms and a fast Chebychev approximation.
result SpGAT achieves better global pattern recognition with fewer parameters compared to GAT.
AAANE embeds networks by learning attention weights for multi-scale structure.
problem Existing methods ignore the role of different scales in network embedding.
method AAANE uses an attention-based adversarial autoencoder to learn robust representations.
result AAANE outperforms existing methods on real-world networks.
Improved speech enhancement with MNTFA using time-frequency attention.
problem Speech enhancement with limited model size and memory.
method Designing MNTFA with self-attention modules for long sequences and joint training.
result MNTFA achieves better performance with fewer parameters than DPCRN.
Replicated attention model in image classification and fine-grained recognition.
problem Improving attention mechanisms in neural networks.
method Implemented the 'Learn to Pay Attention' model in convolutional neural networks.
result Successfully replicated results in image classification and fine-grained recognition.
3D Axial-Attention improves lung nodule classification accuracy.
problem Limited 3D attention in existing methods.
method Proposes 3D Axial-Attention network with 3D positional encoding.
result 3D Axial-Attention achieves state-of-the-art performance.
Lipschitz normalization boosts deep attention models, especially for graph neural networks.
problem Gradient explosion in deep graph attention networks leads to poor performance.
method Enforcing Lipschitz continuity by normalizing attention scores.
result Deep GAT models with LipschitzNorm achieve state-of-the-art results for tasks with long-range dependencies.
Graph Attention Networks use masked self-attention to improve graph neural networks.
problem Improving graph neural networks to better handle graph data.
method Stacked masked self-attention layers that allow nodes to attend to their neighborhoods with different weights.
result GAT models achieve state-of-the-art results across various graph benchmarks.
Graph attention network improves MLTC by capturing label dependencies.
problem Ignoring label dependencies in MLTC tasks.
method Graph attention network model that captures label dependencies.
result The model achieves similar or better performance than state-of-the-art models.
New insights into how encoder-decoder networks generate attention matrices.
problem Understanding how encoder-decoder networks use attention matrices.
method Decomposing hidden states into temporal and input-driven components.
result Attention matrices are formed based on task requirements, not architecture type.
FAN improves attention weights for better relation emphasis.
problem Learning attention weights for better relation emphasis.
method Introduced a novel center-mass cross entropy loss and a focused attention backbone.
result Focused supervision leads to improved attention distribution and enhanced representation.
TSAM predicts directed temporal links using GCN and self-attention.
problem Predicting links in directed temporal networks.
method GCN, self-attention mechanism, autoencoder architecture, graph attentional layers, graph convolutional layers, graph recurrent unit layer.
result TSAM outperforms benchmarks on four realistic networks.
SARN improves relational reasoning with less computation.
problem Efficiently perform relational reasoning with reduced computation.
method Introduces SARN, a sequential attention relational network.
result SARN achieves high accuracy on relational questions.
A new multi-layer attention mechanism improves speech keyword recognition accuracy.
problem Inaccurate attention weights in LSTM networks for speech keyword recognition.
method Introducing information from layers prior to feature extraction into attention weights calculations.
result The proposed multi-layer attention mechanism leads to more accurate attention weights and improved keyword spotting performance.
GSAN learns adaptive node representations using geometric scattering and attention.
problem Oversmoothing in node representation learning.
method Attention-based architecture integrating geometric scattering and GCN channels.
result GSAN outperforms previous networks in semi-supervised node classification.
Proposes a model for multi-agent reinforcement learning with hierarchical graph attention network.
problem Limited transferability of trained policies to new multi-agent tasks.
method Uses hierarchical graph attention network for representation learning and multi-agent actor-critic for policy learning.
result Demonstrates superior performance in mixed cooperative and competitive tasks compared to existing methods.
Models predict emotional valence from narratives, matching human raters.
problem Predicting emotional valence from multimodal time-series data.
method Adapted attention-based mechanisms (Transformer, Memory Fusion Network) to emotional narratives.
result Models perform well, matching human raters on emotional valence prediction.
Bayesian Attention Networks compress data by focusing on key training samples.
problem Lossless data compression for efficiency.
method Bayesian Attention Networks with attention factors and latent space.
result Efficient prediction using a few correlated training samples.
Proposes RN for unsupervised attention in neural networks.
problem Limited, imbalanced, and non-stationary input distributions in various tasks.
method Inspired by neuronal adaptation, RN uses MDL principle and universal code length for incremental layer-wise computation.
result Outperforms existing normalization methods across diverse tasks.
Attention-based GNNs can't prevent oversmoothing, leading to homogeneous node representations.
problem The issue of oversmoothing in attention-based GNNs.
method Viewed attention-based GNNs as nonlinear time-varying dynamical systems and used tools from the theory of products of inhomogeneous matrices and the joint spectral radius.
result Graph attention mechanism cannot prevent oversmoothing and loses expressive power exponentially.
The paper uses attention networks for character-based handwritten text transcription.
problem Handwritten text recognition with improved character-level alignment.
method Attentional encoder-decoder networks trained on character sequences, comparing different activation functions.
result Softmax attention provides more precise character alignment than sigmoid attention.
Graph attention networks improve performance on heterogeneous graphs.
problem Complex performance of GNNs on heterogeneous graphs.
method Integrating positional encodings into graph attention networks.
result Graph attention networks excel in node classification and link prediction.
New hypergraph operators improve graph neural networks for higher-order relationships.
problem Learning deep embeddings on high-order graph-structured data.
method Introducing hypergraph convolution and hypergraph attention operators.
result Extensive experimental results show the effectiveness of hypergraph operators.
ARNe model excels in abstract visual reasoning tasks.
problem Abstract visual reasoning using attention mechanisms.
method Hybrid network architecture combining self-attention and relational reasoning.
result ARNe model surpasses WReN model by 11.28 ppt on PGM datasets.
Enhanced GNN with expanded attention window and partially random embeddings.
problem Limited expressivity of traditional GNNs in distinguishing non-isomorphic graphs.
method Graph attention network with expanding attention window and partially random initial embeddings. Head dropout for regularization.
result Improved ability to differentiate between non-isomorphic graphs.
Variational attention improves latent alignment in NLP tasks.
problem Inefficiency and lack of probabilistic alignment in neural attention.
method Amortized variational inference for latent variable alignment models.
result Variational attention retains performance gains of latent models with comparable training speed.
CRAN extracts music highlights using attention and recurrent layers.
problem Extracting valuable music highlights from signals.
method Convolutional Recurrent Attention Networks (CRAN) with attention mechanism.
result CRAN outperforms three baseline methods in highlighting extraction.
GSA-Nets apply group equivariance to self-attention for vision tasks.
problem Improving self-attention networks for vision tasks.
method Define group-equivariant positional encodings.
result GSA-Nets outperform non-equivariant self-attention networks on vision benchmarks.
AGNN improves network localization accuracy by 37-53% in NLOS conditions.
problem Massive network localization under Non-Line-of-Sight conditions.
method Attentional Graph Neural Network (AGNN) with Adjacency Learning Module (ALM) and Multiple Graph Attention Layers (MGAL).
result Significant improvement in localization accuracy, approaching fundamental lower bounds.
New interpretation of attention in Transformers and Graph Attention Networks.
problem Understanding and improving attention mechanisms in deep learning models.
method Decomposed attention into a kernel and a normalization term; generalized the kernel function and norm.
result Generalized attention leads to better performance on various tasks.
Dual-attention GCN improves text classification by adapting to textual complexity.
problem Challenges in learning discriminative features from texts due to graph variants.
method Proposes a dual-attention GCN with connection-attention and hop-attention mechanisms.
result Achieves state-of-the-art performance on text classification tasks.
Researchers improve transformer networks' optimization and understanding.
problem Improving the understanding and optimization of transformer networks.
method Introducing a convex alternative to the self-attention mechanism and reformulating the training problem as a convex optimization problem.
result Revealed an implicit regularization mechanism that promotes sparsity across tokens.