Efficiently accelerates attention calculation for Transformers with relative positional encoding.
problem Quadratic complexity of attention in long sequences.
method Kernelized attention with Fast Fourier Transform (FFT) for RPE.
result Achieves O(n log n) time complexity, mitigates training instability, and outperforms other models.
Paper introduces a new method for Transformers with linear complexity.
problem No efficient relative positional encoding for linear Transformer models.
method Stochastic Positional Encoding (SPE) that replaces classical RPE.
result SPE behaves like RPE and performs well on benchmarks.
BiPE blends intra-segment and inter-segment encodings for better length extrapolation.
problem Improving length extrapolation in language models.
method Bilevel Positional Encoding (BiPE) that separates intra-segment and inter-segment encodings.
result BiPE enhances length extrapolation across various text modalities.
Randomized positional encodings boost transformer performance on longer sequences.
problem Transformers struggle with generalizing to sequences of arbitrary length.
method Introduced randomized positional encodings that simulate longer sequences and randomly select positions.
result Randomized positional encodings increase test accuracy by 12.0% on average for sequences of unseen length.
GTA improves transformer-based NVS models by encoding geometric structure.
problem Suboptimal positional encoding for 3D vision tasks.
method Geometry-aware attention mechanism encoding geometric structure of tokens.
result GTA improves learning efficiency and performance of NVS models.
Graph attention networks improve performance on heterogeneous graphs.
problem Complex performance of GNNs on heterogeneous graphs.
method Integrating positional encodings into graph attention networks.
result Graph attention networks excel in node classification and link prediction.
We construct a combinatorial invariant of Legendrian knots in standard contact three-space. This invariant, which encodes rational relative Symplectic Field Theory and extends contact homology, counts holomorphic disks with an arbitrary number of positive punctures. The construction uses ideas from string topology.
New analysis shows RPE-based Transformers can't approximate all functions.
problem Understanding the limitations of RPE-based Transformers in approximating continuous functions.
method Mathematical analysis and development of a novel attention module (URPE) to overcome limitations.
result RPE-based Transformers can't approximate all continuous sequence-to-sequence functions, even with depth and width.
ETC improves Transformer models for long and structured inputs.
problem Scaling input length and encoding structured inputs in Transformers.
method Introduces global-local attention, relative position encodings, and CPC pre-training.
result Achieves state-of-the-art results on four natural language datasets.
FraudTransformer detects payment fraud by preserving event order and time gaps.
problem Detecting payment fraud in real-world banking streams with irregular time gaps.
method Augments a GPT-style architecture with a dedicated time encoder and a learned positional encoder.
result FraudTransformer outperforms classical and transformer baselines, achieving highest AUROC and PRAUC on held-out test set.
Introduce Collapsed Effective Operators for higher-order structures.
problem Existing spectral operators decompose topology into separate ranks, leaving practitioners to fuse information back to vertices.
method Introduce Collapsed Effective Operators via Schur complementation of a graded Laplacian.
result Preserves positive semi-definiteness, lowers system energy under higher-order connectivity.
Financial fraud detection in digital banking requires reasoning over multiple heterogeneous event streams.
problem Financial fraud detection in digital banking requires reasoning over multiple heterogeneous event streams.
method Multi-Stream Fraud Transformer (MSFT) architecture that encodes each event stream with independent Transformer encoders and fuses their representations through configurable mechanisms.
result Sequence models significantly outperform gradient-boosted trees operating on aggregated features.
STRING improves 2D and 3D position encodings for better performance.
problem Efficient and accurate position encoding for 2D and 3D applications.
method STRING extends Rotary Position Encodings with a unifying theoretical framework, maintaining translation invariance and low computational cost.
result STRING shows substantial gains in open-vocabulary object detection and robotics.
The paper describes fitting submanifolds to data using Sussmann's orbit theorem.
problem Fitting an immersed submanifold to random samples.
method Uses Sussmann's orbit theorem to ensure submanifold fitting. Reconstruction involves encoding times and decoding via flows of vector fields.
result A high-probability bound on excess risk for the reconstruction error.
New method for Transformer models to encode position information without sequential bias.
problem Lack of flexible and learnable position encoding for Transformer models.
method Continuous dynamical model to learn position encoding.
result Consistent improvements over baselines in various NLP tasks.
A new method, REC, compresses images by encoding their latent representations efficiently.
problem Efficiently compressing single images with latent representations.
method Relative Entropy Coding (REC) that directly encodes latent representations with codelength close to relative entropy.
result REC is more efficient for single image compression compared to previous methods and is competitive for lossy compression.
GraphReach improves GNN performance by incorporating node positions.
problem Existing GNNs fail to capture node positions, leading to inaccurate predictions.
method GraphReach uses reachability estimations from anchor nodes to capture global node positions.
result GraphReach achieves up to 40% relative improvement in accuracy compared to state-of-the-art GNNs.
Transformer learns graph structure better with subgraph info.
problem Transformer struggles with structural similarity in graph learning.
method Structure-Aware Transformer with subgraph attention.
result Improves graph prediction benchmarks significantly.
Paper finds conditions for non-Einstein relative Yamabe metrics.
problem Finding relative Yamabe metrics with positive scalar curvature.
method Sufficient condition for positive constant scalar curvature metrics on manifolds with boundary.
result Examples of non-Einstein relative Yamabe metrics with positive scalar curvature.
Transformer struggles with arithmetic length but improves with explicit structure encoding.
problem Transformers fail to generalize length in arithmetic tasks.
method Explicitly encoding structural symmetries via modified number formatting and custom positional encodings.
result Transformer can generalize up to 50-digit numbers without additional data.
Transformer model improves source code summarization.
problem Generating readable summaries of source code.
method Transformer model with self-attention mechanism for code representation.
result Transformer model outperforms state-of-the-art techniques.
Neural machine translation is a relatively new approach to statistical machine translation based purely on neural networks. The neural machine translation models often consist of an encoder and a decoder. The encoder extracts a fixed-length representation from a variable-length input sentence, and the decoder generates…
This paper studies the prediction of chord progressions for jazz music by relying on machine learning models. The motivation of our study comes from the recent success of neural networks for performing automatic music composition. Although high accuracies are obtained in single-step prediction scenarios, most models fa…
Relative notions of combinatorial asphericity have been used to prove that injective labeled oriented trees (which encode spines of ribbon 2-knots) are aspherical. This article presents an overview and comparison of the different notions of relative combinatorial asphericity. It also contains new results concerning cha…
Paper proposes LCP for structural encodings, outperforming existing methods.
problem Improving Graph Neural Networks performance through effective structural encodings.
method Geometric perspective, Local Curvature Profiles (LCP) for structural encodings, combining with global positional encodings, comparing with rewiring techniques.
result LCP significantly outperforms existing structural encodings and combining LCP with global positional encodings improves performance.
Study a relative aspherical conjecture and prove 3-manifold obstruction to positive scalar curvature.
problem Obstructing the existence of positive scalar curvature in higher dimensions.
method Introduced a relative aspherical condition and a new geometric quantity called spherical width.
result Proved results on how 3-manifolds obstruct the existence of positive scalar curvature.
In this paper, we consider surfaces in 4--dimensional pseudo--Riemannian space--forms with index 2. First, we obtain some of geometrical properties of such surfaces considering their relative null space. Then, we get classifications of quasi--minimal surfaces with positive relative nullity.
The paper studies a relative version of non-positive immersion for 2-complex pairs and shows conditions under which a transitivity law holds.
problem The study of collapsing non-positive immersion for 2-complex pairs and its implications.
method Introduced a relative version of collapsing non-positive immersion for 2-complex pairs (L,K) and proved a transitivity law under certain conditions. result Under certain conditions, a transitivity law holds: If (L,K) has relative collapsing non-positive immersion and K has collapsing non-positive immersion, then L has collapsing non-positive immersion. We define a relative Yamabe invariant of a smooth manifold with given conformal class on its boundary. In the case of empty boundary the invariant coincides with the classic Yamabe invariant. We develop approximation technique which leads to gluing theorems of two manifolds along their boundaries for the relative Yamab…
In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of online attention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decod…
Study proves positivity of quasi-local masses in general relativity using spinors.
problem Proving the positivity of quasi-local masses in general relativity.
method Using spinors and solving Dirac equation on compact Riemannian manifolds with boundary conditions.
result Gravitational mass bounded by a spacelike topological 2-sphere is non-negative, vanishing only in Minkowski space.
Unified framework analyzes and compares RFF and RoPE PEs for music generation.
problem Efficiently modeling music generation with positional encodings.
method Kernel methods to analyze and compare RFF and RoPE PEs.
result RoPEPool outperforms other methods in melody harmonization.
Binary encoding enables neural networks to extrapolate periodic functions.
problem Extrapolating periodic functions without prior knowledge of their form.
method Normalized Base-2 Encoding (NB2E) for continuous numerical values.
result MLPs using NB2E can successfully extrapolate diverse periodic signals.
Article generalizes open book construction for 5D contact pairs.
problem Constructing compatible open books on relative contact pairs.
method Introduces generalized square bridge position for 5D Legendrian links.
result Algorithm constructs relative open book decompositions on relative contact pairs.
We deal with hypersurfaces in the framework of the n-dimensional relative differential geometry. We consider a hypersurface Φ of Rn+1 with position vector field x, which is relatively normalized by a relative normalization y. Then y is also a relative normalizati…
New groups with special properties found.
problem Finding new groups with specific geometric properties.
method Proved actions on CAT(0) cubical complexes under certain conditions.
result Many groups admit cocompact actions on CAT(0) cubical complexes.
We present an attention-based ranking framework for learning to order sentences given a paragraph. Our framework is built on a bidirectional sentence encoder and a self-attention based transformer network to obtain an input order invariant representation of paragraphs. Moreover, it allows seamless training using a vari…
Study space-like and time-like surfaces in Robertson-Walker space-times with positive nullity.
problem Characterize space-like and time-like surfaces in Robertson-Walker space-times with positive relative nullity.
method Provide necessary and sufficient conditions, local classification theorems, and analyze special spaces.
result Local classification theorems for space-like and time-like surfaces in L14(f,0) with positive relative nullity. We discuss some aspects about the computation of kinematic, spectroscopic, Fermi and astrometric relative velocities that are geometrically defined in general relativity. Mainly, we state that kinematic and spectroscopic relative velocities only depend on the 4-velocities of the observer and the test particle, unlike F…
This paper explores GNN functions on random graphs, highlighting the importance of node Positional Encodings.
problem Understanding the expressive power of GNNs on large random graphs.
method General convergence notions, input node features, and Positional Encodings (PEs).
result GNNs can converge to certain functions on large random graphs, emphasizing the role of PEs.
Burau representation of the Artin braid group remains as one of the very important representations for the braid group. Partly, because of its connections to the Alexander polynomial which is one of the first and most useful invariants for knots and links. In the present work, we show that interesting representations o…
New analysis shows PE in Transformers increases generalization gap and vulnerability.
problem Understanding the impact of PE on Transformer generalization and robustness.
method Generalization analysis and adversarial Rademacher bounds for a single-layer Transformer with trainable PE.
result PE systematically enlarges the generalization gap and makes models more vulnerable to attacks.
Study eta invariant on non-compact manifolds with positive scalar curvature.
problem Proving geometric formulas and index theorems for uniformly positive scalar curvature metrics.
method Using Dirac-Schrödinger operators and relative eta invariant.
result New geometric formula for spectral flow and index formula for uniformly positive scalar curvature metrics.
Defines new Roe algebras for cylindrical spaces, solving metric curvature problems.
problem Existence and classification of metrics with positive scalar curvature on spaces with cylindrical ends.
method Variant of Roe algebras for cylindrical spaces, relating to relative higher index theory.
result Defines higher rho-invariants and provides a concise proof of a related result.
This paper is devoted to the 3-dimensional relative differential geometry of surfaces. In the Euclidean space RE3 we consider a surface Φ with position vector field $\vect{x}$, which is relatively normalized by a relative normalization $\vect{y}% (u^1,u^2) $. A surf…
PE-GQNN improves spatial data prediction and uncertainty quantification.
problem Poor calibration of predictive distributions in spatial data models.
method Combines PE-GNNs with Quantile Neural Networks and recalibration techniques.
result PE-GQNN outperforms existing methods in predictive accuracy and uncertainty quantification.
Finsler metrics with relatively non-negative (non-positive, respectively), constant and isotropic stretch curvatures are investigated in this paper. In particular, it is proved that every non-Riemannian (α,β)-metric with a nonzero constant flag curvature and a non-zero relatively isotropic stretch curvature over a m…
GSA-Nets apply group equivariance to self-attention for vision tasks.
problem Improving self-attention networks for vision tasks.
method Define group-equivariant positional encodings.
result GSA-Nets outperform non-equivariant self-attention networks on vision benchmarks.