Attention layers are sensitive to single words, improving generalization over random features.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New invariants derived from random matrices for words in free groups.
We introduce a training method for both better word representation and performance, which we call GROVER (Gradual Rumination On the Vector with maskERs). The method is to gradually and iteratively add random noises to word embeddings while training a model. GROVER first starts from conventional training process, and th…
We present algorithms for topic modeling based on the geometry of cross-document word-frequency patterns. This perspective gains significance under the so called separability condition. This is a condition on existence of novel-words that are unique to each topic. We present a suite of highly efficient algorithms based…
A simple text model shows word lengths follow Zipf's law.
SAFER method certifies robustness to word substitutions without model structure.
RNA structures show that a significant portion of bases do not form hydrogen bonds.
Study on quantitative aspects of trace polynomials in free groups.
Word embedding, which encodes words into vectors, is an important starting point in natural language processing and commonly used in many text-based machine learning tasks. However, in most current word embedding approaches, the similarity in embedding space is not optimized in the learning. In this paper we propose a …
We show that, for any (symmetric) finite generating set of the Torelli group of a closed surface, the probability that a random word is not pseudo-Anosov decays exponentially in terms of the length of the word.
This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context vectors, and text generation. These assumptions are well supported either empirically o…
Word embeddings are ubiquitous in NLP and information retrieval, but it is unclear what they represent when the word is polysemous. Here it is shown that multiple word senses reside in linear superposition within the word embedding and simple sparse coding can recover vectors that approximately capture the senses. The …
We prove a rigidity theorem for the geometry of the unit ball in random subspaces of the scl norm in B_1^H of a free group. In a free group F of rank k, a random word w of length n (conditioned to lie in [F,F]) has scl(w)=log(2k-1)n/6log(n) + o(n/log(n)) with high probability, and the unit ball in a subspace spanned by…
Continuous vector representations of words and objects appear to carry surprisingly rich semantic content. In this paper, we advance both the conceptual and theoretical understanding of word embeddings in three ways. First, we ground embeddings in semantic spaces studied in cognitive-psychometric literature and introdu…
We combine concepts from random matrix theory and free probability together with ideas from the theory of commutator length in groups and maps from surfaces, and establish new connections between the two. More particularly, we study measures induced by free words on the unitary groups . Every word in the free…
We develop necessary and sufficient conditions and a novel provably consistent and efficient algorithm for discovering topics (latent factors) from observations (documents) that are realized from a probabilistic mixture of shared latent factors that have certain properties. Our focus is on the class of topic models in …
New spectral Dehn function characterizes word-hyperbolic groups.
We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
We prove sharp limit theorems on random walks on graphs with values in finite groups. We then apply these results (together with some elementary algebraic geometry, number theory, and representation theory) to finite quotients of lattices in semisimple Lie groups (specifically SL(n,Z) and Sp(2n, Z) to show that a ``ran…
The goal of homomorphic encryption is to encrypt data such that another party can operate on it without being explicitly exposed to the content of the original data. We introduce an idea for a privacy-preserving transformation on natural language data, inspired by homomorphic encryption. Our primary tool is {\em obfusc…
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polyse…
Since the 1970's, physicists and mathematicians who study random matrices in the GUE or GOE models are aware of intriguing connections between integrals of such random matrices and enumeration of graphs on surfaces. We establish a new aspect of this theory: for random matrices sampled from the group $\mathcal{U}\left(n…
The paper introduces a method to improve adversarial robustness in neural networks using randomized perturbations.
In probabilistic approaches to classification and information extraction, one typically builds a statistical model of words under the assumption that future data will exhibit the same regularities as the training data. In many data sets, however, there are scope-limited features whose predictive power is only applicabl…
In this paper we explore the "vector semantics" problem from the perspective of "almost orthogonal" property of high-dimensional random vectors. We show that this intriguing property can be used to "memorize" random vectors by simply adding them, and we provide an efficient probabilistic solution to the set membership …
We extend some properties of random walks on hyperbolic groups to random walks on convergence groups. In particular we prove that if a convergence group acts on a compact metrizable space with the convergence property then we can provide with a compact topology such that random walks on converge a…
Topic modeling analyzes documents to learn meaningful patterns of words. For documents collected in sequence, dynamic topic models capture how these patterns vary over time. We develop the dynamic embedded topic model (D-ETM), a generative model of documents that combines dynamic latent Dirichlet allocation (D-LDA) and…
Locally-contextual CRFs improve sequence labeling performance.
Every pseudo-Anosov mapping class defines an associated veering triangulation of a punctured mapping torus. We show that generically, is not geometric. Here, the word "generic" can be taken either with respect to random walks in mapping class groups or with respect to counting geodesic…
We consider random walks on the mapping class group that have finite first moment with respect to the word metric, whose support generates a non-elementary subgroup and contains a pseudo-Anosov map whose invariant Teichmuller geodesic is in the principal stratum of quadratic differentials. We show that a Teichmuller ge…
Random groups prove length constraints on product of conjugates.
Random groups with high density have Property (T).
The simplicial condition and other stronger conditions that imply it have recently played a central role in developing polynomial time algorithms with provable asymptotic consistency and sample complexity guarantees for topic estimation in separable topic models. Of these algorithms, those that rely solely on the simpl…
An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. The current practice of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential us…
Neural network optimizes learning sequence for reading words.
We introduce a new random group model called the square model: we quotient a free group on generators by a random set of relations, each of which is a reduced word of length four. We prove, as in the Gromov density model, that for densities a random group in the square model is trivial with overwhel…
Chinese word segmentation (CWS) is a fundamental task for Chinese language understanding. Recently, neural network-based models have attained superior performance in solving the in-domain CWS task. Last year, Bidirectional Encoder Representation from Transformers (BERT), a new language representation model, has been pr…
Random hyperbolic surfaces have a spectral gap that approaches 1/4 as genus grows.
Random forests have proven to be reliable predictive algorithms in many application areas. Not much is known, however, about the statistical properties of random forests. Several authors have established conditions under which their predictions are consistent, but these results do not provide practical estimates of ran…
We consider two random group models: the hexagonal model and the square model, defined as the quotient of a free group by a random set of reduced words of length four and six respectively. Our first main result is that in this model there exists a sharp density threshold for Kazhdan's Property (T) and it equals 1/3. Ou…
We study Linial-Meshulam random 2-complexes, which are two-dimensional analogues of Erdős-Rényi random graphs. We find the threshold for simple connectivity to be p = n^{-1/2}. This is in contrast to the threshold for vanishing of the first homology group, which was shown earlier by Linial and Meshulam to be p = 2 log(…
The popular Alternating Least Squares (ALS) algorithm for tensor decomposition is efficient and easy to implement, but often converges to poor local optima---particularly when the weights of the factors are non-uniform. We propose a modification of the ALS approach that is as efficient as standard ALS, but provably rec…
A new whole-sentence language model - neural trans-dimensional random field language model (neural TRF LM), where sentences are modeled as a collection of random fields, and the potential function is defined by a neural network, has been introduced and successfully trained by noise-contrastive estimation (NCE). In this…
We study Linial-Meshulam random 2-complexes, which are two-dimensional analogues of Erdős-Rényi random graphs. We find the threshold for simple connectivity to be p = n^{-1/2}. This is in contrast to the threshold for vanishing of the first homology group, which was shown earlier by Linial and Meshulam to be p = 2 log(…
Given a measure on the Thurston boundary of Teichmueller space, one can pick a geodesic ray joining some basepoint to a randomly chosen point on the boundary. Different choices of measures may yield typical geodesics with different geometric properties. In particular, we consider two families of measures: the ones whic…
Community detection has been an active research area for decades. Among all probabilistic models, Stochastic Block Model has been the most popular one. This paper introduces a novel probabilistic model: RW-HDP, based on random walks and Hierarchical Dirichlet Process, for community extraction. In RW-HDP, random walks c…
Studies across many disciplines have shown that lexical choice can affect audience perception. For example, how users describe themselves in a social media profile can affect their perceived socio-economic status. However, we lack general methods for estimating the causal effect of lexical choice on the perception of a…
This study compares machine learning methods for high-cardinality categorical variables.