Randomized positional encodings boost transformer performance on longer sequences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Non-parametric estimators improve quickest changepoint detection under irregular sequence lengths.
Recurrent neural networks and sequence to sequence models require a predetermined length for prediction output length. Our model addresses this by allowing the network to predict a variable length output in inference. A new loss function with a tailored gradient computation is developed that trades off prediction accur…
The paper examines conditions for compactness in sequences of warped product length spaces.
This research proves guarantees on sequence models' generalization to longer and novel sequences.
We discuss algorithms for estimating the Shannon entropy h of finite symbol sequences with long range correlations. In particular, we consider algorithms which estimate h from the code lengths produced by some compression algorithm. Our interest is in describing their convergence with sequence length, assuming no limit…
The paper offers generalization bounds for Transformers that ignore sequence length.
New bounds show transformers need longer training for length generalization.
Transformers can interpolate finite input sequences exactly.
Despite strong performance on a variety of tasks, neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition. We study the related issue of receiving infinite-length sequences from a recurrent language model when using common decoding algorithm…
New techniques improve channel prediction in noisy wireless systems.
Accelerates signature kernel computation for sequences.
Deep architecture such as hierarchical semi-Markov models is an important class of models for nested sequential data. Current exact inference schemes either cost cubic time in sequence length, or exponential time in model depth. These costs are prohibitive for large-scale problems with arbitrary length and depth. In th…
We will give a weak energy identity for Sacks-Uhlenbeck approximation of harmonic maps and calculate the length of the necks.
J.P. Levine introduced a clover link to investigate the indeterminacy of the Milnor invariants of a link. It is shown that for a clover link, the Milnor numbers of length at most are well-defined if those of length at most vanish, and that the Milnor numbers of length at least are not well-defined if …
Time-based model improves recommendation performance.
New attention model removes softmax, improving sequence length bias.
Labyrinth fractals are self-similar dendrites in the unit square that are defined with the help of a labyrinth set or a labyrinth pattern. In the case when the fractal is generated by a horizontally and vertically blocked pattern, the arc between any two points in the fractal has infinite length [Cristea\&Steinsky 2009…
Study SGD dynamics in sequence models, revealing training phases and influence of sequence length.
The paper provides a theoretical justification for using stable SSM blocks in deep sequential models.
HGConv uses HRR to efficiently detect malware, outperforming existing methods.
Models such as Sequence-to-Sequence and Image-to-Sequence are widely used in real world applications. While the ability of these neural architectures to produce variable-length outputs makes them extremely effective for problems like Machine Translation and Image Captioning, it also leaves them vulnerable to failures o…
New approach reduces malware detection memory requirements and speeds up training.
Phylogenetic tree inference using deep DNA sequencing is reshaping our understanding of rapidly evolving systems, such as the within-host battle between viruses and the immune system. Densely sampled phylogenetic trees can contain special features, including "sampled ancestors" in which we sequence a genotype along wit…
Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length. In this paper we introduce sparse factorizations of the attention matrix which reduce this to . We also introduce a) a variation on architecture and initialization to train deeper net…
Paper analyzes regret bounds for unconstrained online optimization.
This work is devoted to the study of deformations of hyperbolic cone structures under the assumption that the lengths of the singularity remain uniformly bounded over the deformation. Given a sequence of pointed hyperbolic cone-manifolds with topological type , where is a closed, orientab…
In this paper we study 1/k geodesics, those closed geodesics that minimize on all subintervals of length , where is the length of the geodesic. We develop new techniques to study the minimizing properties of these curves on doubled polygons, and demonstrate a sequence of doubled polygons whose closed geodesics…
TOQ-Nets learn to recognize complex temporal events with varying objects and sequences.
While neural sequence generation models achieve initial success for many NLP applications, the canonical decoding procedure with left-to-right generation order (i.e., autoregressive) in one-pass can not reflect the true nature of human revising a sentence to obtain a refined result. In this work, we propose XL-Editor, …
We present the Latent Sequence Decompositions (LSD) framework. LSD decomposes sequences with variable lengthed output units as a function of both the input sequence and the output sequence. We present a training algorithm which samples valid extensions and an approximate decoding algorithm. We experiment with the Wall …
Representation learning of pedestrian trajectories transforms variable-length timestamp-coordinate tuples of a trajectory into a fixed-length vector representation that summarizes spatiotemporal characteristics. It is a crucial technique to connect feature-based data mining with trajectory data. Trajectory representati…
BigBird improves transformer performance on NLP tasks with longer sequences.
This paper explores using a Long short-term memory (LSTM) based sequence autoencoder to learn interesting features for detecting surveillance aircraft using ADS-B flight data. An aircraft periodically broadcasts ADS-B (Automatic Dependent Surveillance - Broadcast) data to ground receivers. The ability of LSTM networks …
Improved phylogenetic tree reconstruction using flexible branch length distributions.
For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of Mel-spectrogram augmentation on training the sequence-to-sequence voice conversion (VC…
We study the asymptotic behavior of the asymptotic translation lengths on the curve complexes of pseudo-Anosov monodromies in a fibered cone of a fibered hyperbolic 3-manifold with . For a sequence of fibers and monodromies in the fibered cone, we show that the asymptotic translation len…
New language model shows context length impacts generation quality and reasoning ability.
This study examines how sequential correlations affect in-context learning in sequence models.
Task hinting improves transformer performance on longer tasks.
Latent Block-Diffusion Temporal Point Processes (LBDTPP) is a semi-autoregressive framework for generating asynchronous event sequences.
An online framework optimizes efficiency in conformal prediction with a target miscoverage rate.
Stratifying patients at risk for postoperative complications may facilitate timely and accurate workups and reduce the burden of adverse events on patients and the health system. Currently, a widely-used surgical risk calculator created by the American College of Surgeons, NSQIP, uses 21 preoperative covariates to asse…
Clustered attention improves transformer efficiency for large sequences.
Proposes using frequent sequences to improve sequential recommendation models.
Sum-Product Networks (SPN) have recently emerged as a new class of tractable probabilistic graphical models. Unlike Bayesian networks and Markov networks where inference may be exponential in the size of the network, inference in SPNs is in time linear in the size of the network. Since SPNs represent distributions over…
Music relies heavily on repetition to build structure and meaning. Self-reference occurs on multiple timescales, from motifs to phrases to reusing of entire sections of music, such as in pieces with ABA structure. The Transformer (Vaswani et al., 2017), a sequence model based on self-attention, has achieved compelling …
Transformer model predicts stock trends using technical data and sentiment analysis.