Paper analyzes infinite-width attention layers using Tensor Programs.
problem Capturing the infinite-width limit of attention layers.
method Tensor Programs framework to rigorously identify the limit distribution.
result Derives exact form of infinite-width limit distribution without Gaussian approximations.
This work analyzes self-attention matrices using random matrix theory.
problem Understanding the theoretical behavior of self-attention layers in neural networks.
method Asymptotic spectral analysis of the attention matrix, Gaussian equivalence, and linearization.
result The singular value distribution of the attention matrix is asymptotically characterized by a linear model.
Proposes GPCA module for channel attention in CNNs using Gaussian processes.
problem Improving performance in visual tasks through effective channel selection.
method Integrates Gaussian processes into channel attention mechanisms for probabilistic modeling of channel correlations.
result Demonstrates improved performance of GPCA module in end-to-end CNN training.
Attention mechanisms in deep learning become Gaussian process-like as the number of heads increases.
problem Understanding the behavior of attention mechanisms in deep learning models.
method Extending the equivalence between wide neural networks and Gaussian processes to attention architectures.
result Multi-head attention architectures behave as Gaussian processes as the number of heads tends to infinity.
Scales attention for long contexts in LLMs.
problem Development of attention mechanisms for long context inference.
method Scale-invariant total attention and sparsity conditions, with a position-dependent transformation of logits.
result Scale-invariant attention scheme improves validation loss and long-context retrieval.
Proposes a neural network model for embedding knowledge bases and answering questions.
problem Handling uncertainty and conjunction in neural question answering.
method Gaussian attention model for neural memory access and scoring function.
result Demonstrates model's effectiveness on soccer player dataset for path and conjunctive queries.
A new method for self-attention models that improves uncertainty estimation.
problem Overconfident predictions and lack of calibrated uncertainty in Transformers.
method Kernel-Eigen Pair Sparse Variational Gaussian Processes (KEP-SVGP) with Kernel SVD (KSVD) to handle asymmetry of attention kernels.
result Reduction in time complexity and improved performance on various benchmarks.
A new model uses attention and Gaussian processes for efficient time-series generation.
problem Computational inefficiency and uncertainty underestimation in sequence transduction.
method Attention-based Gaussian process network for real-valued sequence generation.
result The model improves training efficiency and learns factorized generative distribution.
Transformers can solve complex filtering problems for non-Gaussian signals.
problem Non-linear and non-Markovian filtering problems for conditionally Gaussian signals.
method Continuous-time transformer models called filterformers.
result Filterformers can approximate the conditional law of non-Markovian and conditionally Gaussian signal processes.
Skyformer uses Gaussian kernel and Nyström method to speed up self-attention in transformers.
problem High computational cost of self-attention in transformers.
method Replaces softmax with Gaussian kernel and applies Nyström method for matrix approximation.
result Skyformer achieves comparable or better performance with fewer computation resources.
Generative models use kernel smoothing for conditioning on small example sets.
problem Improving generative models' performance with limited conditioning examples.
method Showed that cross-attention conditioning is equivalent to kernel smoothing, specifically a Nadaraya--Watson kernel smoother.
result The approach predicts and confirms three failure regimes for kernel-based conditioning.
Transformers can cluster data from Gaussian mixtures without supervision.
problem Clustering data from Gaussian mixtures without labeled data.
method Theoretical analysis of attention-based layers, focusing on a simplified two-head attention layer and an identity matrix attention layer.
result Attention-based layers can align with true mixture centroids and adapt to input-specific distributions.
Transformer-MGK replaces redundant heads with Gaussian key mixtures, improving efficiency and performance.
problem Redundant attention heads in transformers degrade performance and efficiency.
method Transformer-MGK replaces redundant heads with a mixture of Gaussian keys.
result Transformer-MGK accelerates training and inference, reduces parameters and FLOPs, and achieves comparable or better accuracy.
Attention learns PCA on Gaussian data, proving its connection to principal component analysis.
problem Principal component analysis on Gaussian data.
method Analysis of attention mechanisms through PCA, covering finite and infinite prompt regimes.
result Attention aligns with principal eigenvectors of covariance matrices, converging to optimal solutions in the infinite-prompt limit.
Set Transformer tackles permutation-invariant tasks using attention mechanisms.
problem Tasks defined on sets of instances where order does not matter.
method Attention-based neural network module with an encoder and decoder using attention mechanisms. Introduced a new attention scheme to reduce computational complexity.
result Demonstrates state-of-the-art performance on various set-structured tasks.
Unified framework for studying softmax attention under large prompts.
problem Challenges in theoretical analysis of softmax attention.
method Measure-based framework for finite and infinite prompts.
result Softmax attention converges to linear attention in the large-prompt regime.
Transformers improve with Fourier integral attentions.
problem Inefficiency of dot-product attention in capturing feature dependencies.
method Interpreted attention as kernel regression, proposed FourierFormer with generalized Fourier integral kernels.
result FourierFormer achieves better accuracy and reduces redundancy.
SGPA calibrates transformer uncertainty for safety-critical tasks.
problem Uncertainty estimation in transformer models for safety-critical domains.
method Bayesian inference in transformer's output space using sparse Gaussian processes.
result SGPA-based Transformers improve in-distribution calibration and out-of-distribution robustness.
Graph attention improves node classification by distinguishing important edges.
problem Node classification in graph-based learning models.
method Theoretical analysis of graph attention networks for node classification.
result Graph attention can perfectly classify nodes in an 'easy' regime but fails in a 'hard' regime.
Improved Transformer performance by addressing 'explaining away' effect.
problem Transformer's self-attention mechanism can explain away important input features.
method Proposed a doubly-normalized attention scheme to avoid 'explaining away' effect.
result Improved performance on benchmarks with the new attention scheme.
Fast Bayesian inference with adaptable priors for real-time applications.
problem Intractable exact posterior computation limits Bayesian inference's adoption.
method Distribution Transformer architecture that learns mappings between priors and posteriors.
result Significant reduction in computation time from minutes to milliseconds.
Researchers develop a new spatial process model for non-Gaussian data.
problem Non-Gaussian spatial data with asymmetry and heavy-tailedness.
method Re-parameterized Unified Skew-Normal (SUN) distribution, GSUN process, neural Bayes inference with GATs.
result GSUN process captures non-Gaussian spatial data properties and outperforms conventional models.
Paper introduces kernel deformed exponential families for sparse continuous attention.
problem Creating efficient attention mechanisms for sparse data.
method Developed kernel deformed exponential families, theoretically and experimentally.
result Kernel deformed exponential families can attend to multiple compact regions of data.
This paper develops sparse alternatives to continuous distributions, including new types of Gaussians and attention mechanisms.
problem Creating flexible continuous distributions with varying support for machine learning applications.
method Defining Ω-regularized prediction maps and Fenchel-Young losses for arbitrary domains, and deriving new types of Gaussians and attention mechanisms. result Sparse alternatives to continuous distributions, including deformed exponential families and β-Gaussians, are introduced. NGD improves multivariate Gaussian inference by optimizing Fisher information.
problem Efficiently optimizing multivariate Gaussian models.
method Natural Gradient Descent applied to multivariate Gaussian parameters.
result NGD updates are more efficient for symmetric covariance matrices.
Gated attention improves model curvature, enhancing performance on nonlinear tasks.
problem Understanding the geometric implications of gating in attention mechanisms.
method Modeling attention outputs as Gaussian distributions and analyzing Fisher--Rao geometry.
result Gated attention enables non-flat geometries, including positively curved manifolds.
Paper develops Gaussian process for distributions using Wasserstein distances.
problem Forecasting Gaussian processes indexed by probability distributions.
method Developed positive definite kernels based on Wasserstein distances.
result Efficient forecasting of Gaussian processes indexed by distributions.
The paper examines how SGD noise deviates from Gaussian distribution.
problem Understanding why SGD outperforms GD in neural networks.
method Analysis of SGN vectors' distribution during training.
result For large batch sizes, SGN vectors are mostly Gaussian in early phases.
New acquisition function for extreme rewards in bandits.
problem Online decision making with extreme payoffs in multi-armed bandits.
method Modeling payoffs as Gaussian processes and using a novel UCB acquisition function.
result Demonstrated benefits across synthetic and real-world benchmarks.
Transformers learn to cluster Gaussian mixtures as well as the EM algorithm.
problem Learning guarantees of Transformers in multi-class clustering of Gaussian mixtures.
method Developed a theory connecting Transformer's Softmax Attention layers to the EM algorithm's workflow.
result Transformers achieve minimax optimal rate for clustering Gaussian mixtures with sufficient training samples and initialization.
HRFs adaptively linearize kernels for accurate approximations.
problem Linearizing softmax and Gaussian kernels for machine learning applications.
method Generalizes Bochner's Theorem for kernels, uses random features for compositional kernels.
result Strong theoretical guarantees and unbiased approximation with smaller relative errors.
New RFs reduce kernel approximation variance and improve Transformer performance.
problem Efficient approximation of Gaussian and softmax kernels for kernel methods and Transformers.
method Parameterized, positive, non-trigonometric RFs optimized for variance reduction.
result Significant variance reduction in practice, outperforming previous methods.
Linear Transformer Block combines MLP and linear attention for near-optimal ICL in linear regression.
problem Achieving near-optimal in-context learning (ICL) risk for linear regression with a Gaussian prior.
method Combines linear attention and MLP components in a Linear Transformer Block (LTB). Establishes correspondence with one-step gradient descent estimators (GDext−β). result LTB achieves nearly Bayes optimal ICL risk for linear regression with a Gaussian prior.
Paper extends sparse alternatives to softmax for continuous domains, enabling efficient attention mechanisms.
problem Efficiently assigning zero probability to irrelevant categories in continuous domains.
method Extend alpha-entmax to continuous domains, introducing continuous-domain attention mechanisms.
result Continuous attention allows attending to time intervals and compact regions, improving text classification, machine translation, and visual question answering.
Nearly all Gaussian points in high dimensions lie on a common ellipsoid.
problem Finding an ellipsoid that fits a large set of Gaussian points in high dimensions.
method Analyzing a random set of Gaussian points and proving a bound on their concentration.
result The bound nearly confirms a conjecture about fitting Gaussian points to ellipsoids.
This work studies learning a multi-head attention layer from random examples.
problem Learning a multi-head attention layer from random examples.
method The work initiates the study of provably learning a multi-head attention layer from random examples, providing upper and lower bounds.
result The first nontrivial upper and lower bounds for learning a multi-head attention layer from random examples are given.
Diffusion Transformer captures spatial-temporal dependencies in sequential data.
problem Capturing rich spatial and temporal dependencies in sequential data.
method Established theoretical guarantees for diffusion transformers learning Gaussian process data.
result Spatial-temporal dependencies are captured within attention layers of diffusion transformers.
Unified perspective on Hopfield networks with attention module.
problem Understanding and optimizing Hopfield networks with attention mechanisms.
method Study of BM counterparts of modern Hopfield networks and their salient properties.
result Introduction of AttnBM with tractable likelihood and gradient.
Proposes a novel node embedding framework for graphs using Fisher Information.
problem Lack of theoretical understanding of attention-based GNNs.
method Uses hierarchical kernels and Fisher Information to learn node embeddings.
result Proposed method outperforms existing GNNs on node classification benchmarks.
HKT improves sequence processing with multi-scale attention and kernel analysis.
problem Processing sequences at multiple scales with efficient attention mechanisms.
method Trainable causal downsampling and convex weights for level-specific score matrices.
result HKT achieves consistent gains over standard attention across various tasks.
A study on optimizing self-attention in tabular data using Optimal Transport.
problem Improving efficiency and accuracy of self-attention in tabular classification tasks.
method Developed an OT-based algorithm to generate class-specific dummy Gaussian distributions and train an MLP.
result Achieved comparable accuracy to Transformers with reduced computational cost and efficiency.
New Stein identity for q-Gaussians reduces gradient variance in machine learning.
problem Improving gradient estimators for non-Gaussian distributions.
method Deriving a new Stein identity for bounded-support q-Gaussians and simplifying previous results.
result Gradient estimators for q-Gaussians have nearly identical forms to Gaussian ones, reducing variance.
TNP-KR improves scalability of NPs with Transformer blocks and attention mechanisms.
problem Scalability bottleneck in NPs and GPs, especially for large datasets.
method Introduces TNP-KR with KRBlock, kernel-based attention, and two attention mechanisms.
result TNP-KR with DKA outperforms Performer and achieves state-of-the-art results.
Paper connects tensor regression and Gaussian processes for multi-way data analysis.
problem Learning high-order correlations from multi-way data.
method Demonstrates connections between low-rank tensor regression and Gaussian processes, proving oracle inequality and learning curve.
result Low-rank tensor regression is equivalent to constrained Bayesian inference in Gaussian processes, with learning dependent on eigenvalues and variable correlations.
Approaches KL divergence for learning multi-sense word distributions.
problem Capturing the polysemy and uncertainty of words in word embeddings.
method Modeling words as multi-sense Gaussian mixtures and using KL divergence for learning.
result The proposed approach effectively captures word entailment and distribution similarity.
This study examines a single attention layer's capabilities using random features.
problem Understanding the learning and generalization of a single multi-head attention layer.
method Random feature setting with large number of heads, frozen query and key matrices, and trainable value matrices.
result Random-feature attention layer can express a broad class of permutation-invariant target functions.
The paper extends Gaussian processes to model complex interactions in cellular complexes.
problem Capturing topological inductive biases in machine learning models.
method Proposes Gaussian processes on cellular complexes, introducing novel kernels.
result Derives two novel kernels for modeling interactions between cells.
KITT uses transformers to quickly recommend kernels for GP models.
problem Kernel selection for high-dimensional GP regression models.
method Transformer-based architecture for generating kernel recommendations.
result KITT selects kernels that perform well on various regression benchmarks.