Groups with specific curvature have a regular language of geodesics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Neural language modeling (LM) has led to significant improvements in several applications, including Automatic Speech Recognition. However, they typically require large amounts of training data, which is not available for many domains and languages. In this study, we propose a multilingual neural language model archite…
In natural language processing, it has been observed recently that generalization could be greatly improved by finetuning a large-scale language model pretrained on a large unlabeled corpus. Despite its recent success and wide adoption, finetuning a large pretrained language model on a downstream task is prone to degen…
Let be a finitely generated group. We show that for any finite generating set , the language consisting of all geodesics in with a contracting property is a regular language. As an application, we show that any finitely generated group containing an infinite contracting geodesic must be either virtual…
Let be the fundamental group of a manifold modeled on three dimensional Sol geometry. We prove that has a finite index subgroup which has a rational growth series with respect to a natural generating set. We do this by enumerating by a regular language. However, in contrast to most earlier proofs of thi…
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
New principles needed for scaling large language models, challenging traditional regularization methods.
We consider the quantifier-free languages, Bc and Bc0, obtained by augmenting the signature of Boolean algebras with a unary predicate representing, respectively, the property of being connected, and the property of having a connected interior. These languages are interpreted over the regular closed sets of n-dimension…
Recurrent neural networks are a widely used class of neural architectures. They have, however, two shortcomings. First, it is difficult to understand what exactly they learn. Second, they tend to work poorly on sequences requiring long-term memorization, despite having this capacity in principle. We aim to address both…
AFTER technique improves NLP models by preventing overfitting to task-specific domains.
This study analyzes how one-layer transformers learn regular language recognition tasks.
Enhances reward specification in RL with a novel language-based approach.
We prove that the exponential growth rate of the regular language of penetration sequences is smaller than the growth rate of the regular language of normal form words, if the acceptor of the regular language of normal form words is strongly connected. Moreover, we show that the latter property is satisfied for all irr…
Generative Adversarial Networks (GANs) enjoy great success at image generation, but have proven difficult to train in the domain of natural language. Challenges with gradient estimation, optimization instability, and mode collapse have lead practitioners to resort to maximum likelihood pre-training, followed by small a…
New RLHF approach mitigates bias in aligning LLMs with human preferences.
Mitigates gender bias amplification in model predictions.
There has been an increased interest in multimodal language processing including multimodal dialog, question answering, sentiment analysis, and speech recognition. However, naturally occurring multimodal data is often imperfect as a result of imperfect modalities, missing entries or noise corruption. To address these c…
Researchers calculate complexity of billiard paths in regular polygons.
Recurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects random noise into t…
Paper analyzes and improves KL-regularized RL for LLMs with logarithmic regret.
We recast the Calabi flow in DeGiorgi's language of minimizing movements. We establish the long time existence of minimizing movements for K-energy with arbitrary initial condition. Furthermore we establish some a priori regularity of these solutions, and that sufficiently regular minimizing movements are smooth soluti…
AER dynamically adjusts entropy regularization for better LLM reinforcement learning.
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
Cross-language learning allows us to use training data from one language to build models for a different language. Many approaches to bilingual learning require that we have word-level alignment of sentences from parallel corpora. In this work we explore the use of autoencoder-based methods for cross-language learning …
The language of maximal lexicographic representatives of elements in the positive braid monoid with generators is a regular language. We describe with great detail the smallest Finite State Automaton accepting such language, and study the proportion of elements of length whose maximal lexicographic repres…
Paper uses 'manifold attack' to improve model accuracy and robustness.
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
New bounds show large language models can generalize beyond training data.
Generative Adversarial Networks (GANs) have been promising in the field of image generation, however, they have been hard to train for language generation. GANs were originally designed to output differentiable values, so discrete language generation is challenging for them which causes high levels of instability in tr…
AlphaPruning optimizes LLM pruning using HT-SR theory for better performance.
Information extraction is an important task in NLP, enabling the automatic extraction of data for relational database filling. Historically, research and data was produced for English text, followed in subsequent years by datasets in Arabic, Chinese (ACE/OntoNotes), Dutch, Spanish, German (CoNLL evaluations), and many …
New algorithm avoids spurious sharpness minimization for NLP models.
Soft diamond regularizers improve deep learning performance and sparsity.
Kronecker Products (KP) have been used to compress IoT RNN Applications by 15-38x compression factors, achieving better results than traditional compression methods. However when KP is applied to large Natural Language Processing tasks, it leads to significant accuracy loss (approx 26%). This paper proposes a way to re…
A new memory-efficient sign language translation model reduces weight usage.
ATD measures language distance using neural models, recovering linguistic groupings.
Recently, substantial progress has been made in language modeling by using deep neural networks. However, in practice, large scale neural language models have been shown to be prone to overfitting. In this paper, we present a simple yet highly effective adversarial training mechanism for regularizing neural language mo…
Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a large amount of computation and making them prone t…
Efficient echo state network with explicit memory performs well on benchmark tasks.
This paper addresses the problem of inferring a regular expression from a given set of strings that resembles, as closely as possible, the regular expression that a human expert would have written to identify the language. This is motivated by our goal of automating the task of postmasters of an email service who use r…
We survey - by means of 20 examples - the concept of varifold, as generalised submanifold, with emphasis on regularity of integral varifolds with mean curvature, while keeping prerequisites to a minimum. Integral varifolds are the natural language for studying the variational theory of the area integrand if one conside…
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA) method, but only on the subsequence of training examples where a classification error is made. We derive a bound on the number of mistakes…
Language evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural network architectures rather than persons. This type of temporal signal is typically overlooked, but is important if one aims to deploy a mac…
ETC improves Transformer models for long and structured inputs.
Efforts to understand the generalization mystery in deep learning have led to the belief that gradient-based optimization induces a form of implicit regularization, a bias towards models of low "complexity." We study the implicit regularization of gradient descent over deep linear neural networks for matrix completion …
We show that power-law analyses of financial commentaries from newspaper web-sites can be used to identify stock market bubbles, supplementing traditional volatility analyses. Using a four-year corpus of 17,713 online, finance-related articles (10M+ words) from the Financial Times, the New York Times, and the BBC, we s…
We propose a generalization of neural network sequence models. Instead of predicting one symbol at a time, our multi-scale model makes predictions over multiple, potentially overlapping multi-symbol tokens. A variation of the byte-pair encoding (BPE) compression algorithm is used to learn the dictionary of tokens that …
Unified framework for ICL in causal and masked models.