New combinatorial framework for geometric realizations of subword complexes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Byte-level machine translation outperforms embedding-based methods.
The study proves a theorem about subword complexity for free group automorphisms.
We propose several ways of reusing subword embeddings and other weights in subword-aware neural language models. The proposed techniques do not benefit a competitive character-aware model, but some of them improve the performance of syllable- and morpheme-aware models while showing significant reductions in model sizes…
Semantic representations of words have been successfully extracted from unlabeled corpuses using neural network models like word2vec. These representations are generally high quality and are computationally inexpensive to train, making them popular. However, these approaches generally fail to approximate out of vocabul…
A model corrects Lithuanian grammatical errors.
Universal Language Model for Fine-tuning [arXiv:1801.06146] (ULMFiT) is one of the first NLP methods for efficient inductive transfer learning. Unsupervised pretraining results in improvements on many NLP tasks for English. In this paper, we describe a new method that uses subword tokenization to adapt ULMFiT to langua…
This paper undertakes a study of the structure of the fibers of the Chevalley exponentiation maps . The fibers of these maps encode the nonnegative real relations amongst exponentiated Chevalley generators. Our main theorems show that the fibers admit cell stratifications, t…
Study on representations of four-punctured sphere group in hyperbolic spaces.
Sequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition. In this work, we show that such models can achieve competitive results on the Switchboard 300h and LibriSpeech 1000h tasks. In particular, we report the state-of-the-art word error rates (WER) of 3.5…
AV-ASR system improves speech recognition with visual context.
Traditionally, many text-mining tasks treat individual word-tokens as the finest meaningful semantic granularity. However, in many languages and specialized corpora, words are composed by concatenating semantically meaningful subword structures. Word-level analysis cannot leverage the semantic information present in su…
We investigate intersections of geodesic lines in and in an associated tree T, proving the following result. Let M be a punctured hyperbolic torus and let be a closed geodesic in M. Any edge of any triangle formed by distinct geodesic lines in the preimage of in is shorter then . However, a simil…
Acoustic Neighbor Embeddings map speech and text to fixed dimensions for phonetic confusability.
We consider the problem of making machine translation more robust to character-level variation at the source side, such as typos. Existing methods achieve greater coverage by applying subword models such as byte-pair encoding (BPE) and character-level encoders, but these methods are highly sensitive to spelling mistake…
Text normalization is an important enabling technology for several NLP tasks. Recently, neural-network-based approaches have outperformed well-established models in this task. However, in languages other than English, there has been little exploration in this direction. Both the scarcity of annotated data and the compl…
We introduce Probabilistic FastText, a new model for word embeddings that can capture multiple word senses, sub-word structure, and uncertainty information. In particular, we represent each word with a Gaussian mixture density, where the mean of a mixture component is given by the sum of n-grams. This representation al…
Paper reduces vocabulary losslessly for language model cooperation.
SAFER method certifies robustness to word substitutions without model structure.