Regular languages describe contracting geodesics in groups.
problem Characterizing groups with infinite contracting geodesics.
method Analyzing geodesics in Cayley graphs with a contracting property.
result Groups with infinite contracting geodesics are either virtually Z or acylindrically hyperbolic.
Groups with specific curvature have a regular language of geodesics.
problem Understanding the language of geodesics in non-positively curved triangle groups.
method Proving finitely many cone types and regularity of geodesic languages.
result The language of lexicographically first geodesics is regular and satisfies the fellow traveller property.
Multilingual LM improves low-resource language modeling.
problem Lack of data for many languages and domains.
method Jointly trained multilingual neural language model with shared parameters.
result Significant improvements in conversational data domain with limited training data.
Mixout technique improves finetuning of large pretrained models on few instances.
problem Degenerate performance in finetuning large pretrained models with limited training data.
method Mixout technique, inspired by dropout, stochastically mixes model parameters.
result Stability and accuracy of finetuning improve significantly with mixout.
Researchers found a regular language for maximal lexicographic representatives in braid monoids.
problem Understanding the language of maximal lexicographic representatives in braid monoids.
method Detailed description of the smallest Finite State Automaton and analysis of the proportion of elements.
result The proportion of elements of length k whose maximal lexicographic representative finishes with the first generator tends to a number P_{n,1} ≥ 1/8 as k tends to infinity.
Let Γ be the fundamental group of a manifold modeled on three dimensional Sol geometry. We prove that Γ has a finite index subgroup G which has a rational growth series with respect to a natural generating set. We do this by enumerating G by a regular language. However, in contrast to most earlier proofs of thi…
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
problem Stability of internal states in recurrent neural networks trained on regular languages.
method Empirical study with analysis of network activation and transitions between states.
result Recurrent neural networks trained on regular languages can recover from random perturbations and maintain stable states.
New principles needed for scaling large language models, challenging traditional regularization methods.
problem The shift from generalization to scaling in machine learning requires new guiding principles.
method Examining the effectiveness of traditional regularization methods in the scaling-centric era.
result Traditional principles of regularization may not generalize to larger scales, highlighting new phenomena like scaling law crossover.
We consider the quantifier-free languages, Bc and Bc0, obtained by augmenting the signature of Boolean algebras with a unary predicate representing, respectively, the property of being connected, and the property of having a connected interior. These languages are interpreted over the regular closed sets of n-dimension…
Improved neural language models using adversarial training.
problem Overfitting in large-scale neural language models.
method Adversarial training mechanism to introduce adversarial noise to output embedding layers.
result Improved perplexity scores and BLEU scores on various language tasks.
State-regularized RNNs improve interpretability and performance on long-term memory tasks.
problem RNNs struggle with long-term memory and lack of interpretability.
method Introduce a stochastic state transition mechanism to limit state transitions to a finite set.
result State-regularized RNNs perform better on tasks requiring long-term memory.
Model learns tensor representations from imperfect multimodal data.
problem Learning from imperfect multimodal data with noise or missing entries.
method Tensor rank minimization to regularize rank of tensor representations.
result Model effectively learns tensor representations from imperfect data.
AFTER technique improves NLP models by preventing overfitting to task-specific domains.
problem Standard fine-tuning degrades pretraining domain representations.
method Complements task-specific loss with adversarial objective.
result AFTER leads to improved performance on various NLP tasks.
This study analyzes how one-layer transformers learn regular language recognition tasks.
problem Understanding how one-layer transformers solve regular language recognition tasks like even pairs and parity check.
method Theoretical analysis of training dynamics and gradient descent for a one-layer transformer.
result A one-layer transformer can solve even pairs directly but needs CoT for parity check. Training phases show rapid growth in attention layer followed by logarithmic growth in linear layer.
Enhances reward specification in RL with a novel language-based approach.
problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.
We prove that the exponential growth rate of the regular language of penetration sequences is smaller than the growth rate of the regular language of normal form words, if the acceptor of the regular language of normal form words is strongly connected. Moreover, we show that the latter property is satisfied for all irr…
ScratchGAN trains language GANs without pre-training, achieving comparable results.
problem Training GANs for natural language is challenging due to gradient estimation, optimization instability, and mode collapse.
method Used large batch sizes, dense rewards, and discriminator regularization to stabilize and improve language GANs.
result ScratchGAN performs comparably to maximum likelihood training on quality and diversity metrics.
A new regularizer boosts long-range dependency in sequence data.
problem Improving long-range dependency in sequence data models.
method Developed a mutual information regularizer to enhance sequence learning.
result The approach increases mutual information and likelihood on holdout data.
New RLHF approach mitigates bias in aligning LLMs with human preferences.
problem Algorithmic bias in RLHF leading to preference collapse.
method Preference Matching (PM) RLHF, using PM regularizer and conditional variant.
result 29% to 41% improvement in alignment with human preferences.
Mitigates gender bias amplification in model predictions.
problem Gender bias amplification in model predictions.
method Posterior regularization to mitigate bias.
result Almost removes gender bias amplification in model predictions.
Researchers calculate complexity of billiard paths in regular polygons.
problem Calculating the complexity of billiard paths in regular polygons.
method Counting saddle connections on lattice surfaces, focusing on combinatorial length.
result They answered a question about billiard language complexity in regular polygons.
Recurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects random noise into t…
Paper analyzes and improves KL-regularized RL for LLMs with logarithmic regret.
problem Improving efficiency of RL fine-tuning for large language models.
method Optimism-based KL-regularized online contextual bandit algorithm with novel regret analysis.
result Achieves an O(ηlog(NRT)⋅dR) logarithmic regret bound. AER dynamically adjusts entropy regularization for better LLM reinforcement learning.
problem Policy entropy collapse in RLVR training limits exploration and reasoning performance.
method Adaptive Entropy Regularization (AER) with difficulty-aware coefficient allocation, initial-anchored target entropy, and dynamic global coefficient adjustment.
result AER consistently outperforms baselines on mathematical reasoning benchmarks, improving both accuracy and exploration.
We recast the Calabi flow in DeGiorgi's language of minimizing movements. We establish the long time existence of minimizing movements for K-energy with arbitrary initial condition. Furthermore we establish some a priori regularity of these solutions, and that sufficiently regular minimizing movements are smooth soluti…
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.
Cross-language learning allows us to use training data from one language to build models for a different language. Many approaches to bilingual learning require that we have word-level alignment of sentences from parallel corpora. In this work we explore the use of autoencoder-based methods for cross-language learning …
Paper uses 'manifold attack' to improve model accuracy and robustness.
problem Improving model performance with limited training data and large model parameters.
method Enforces manifold preservation from original data into latent space using adversarial learning.
result Improves accuracy rate and robustness to adversarial examples.
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
Cyclical Annealing Schedule improves VAE performance in NLP tasks.
problem KL term vanishing in VAEs with auto-regressive decoders.
method Cyclical Annealing Schedule to vary β over multiple cycles.
result Progressive learning of meaningful latent codes improves VAE performance.
This paper improves compression of large NLP models using doped Kronecker Products.
problem Accuracy loss when compressing large NLP tasks with Kronecker Products.
method Doping Kronecker Products with an overlay matrix to recover accuracy, and a new regularization scheme called co matrix dropout regularization (CMR).
result Compression of a large language model with LSTM layers of size 25 MB by 25x with 1.4% loss in perplexity score.
New bounds show large language models can generalize beyond training data.
problem Generalization of large language models beyond training data.
method Compression bound derivation using prediction smoothing and SubLoRA.
result Large language models can discover generalizable regularities.
Structured dropout reduces transformer depth for efficient training and inference.
problem Overparameterized transformer networks overfit and require large computation.
method LayerDrop structured dropout for efficient pruning and regularization.
result Sub-networks of any depth can be selected from a large network without performance loss.
Generative Adversarial Networks (GANs) have been promising in the field of image generation, however, they have been hard to train for language generation. GANs were originally designed to output differentiable values, so discrete language generation is challenging for them which causes high levels of instability in tr…
AlphaPruning optimizes LLM pruning using HT-SR theory for better performance.
problem Improving pruning of large language models to reduce size without sacrificing performance.
method AlphaPruning uses HT-SR theory to allocate layerwise sparsity ratios more theoretically.
result AlphaPruning prunes LLaMA-7B to 80% sparsity with reasonable perplexity.
Unified BERT model improves NER across multiple languages.
problem Language-specific NER models limit data extraction.
method Jointly trained multilingual BERT with regularization.
result Unified model outperforms monolingual models on various datasets.
New algorithm avoids spurious sharpness minimization for NLP models.
problem SAM fails in NLP, leading to performance degradation.
method Developed Functional-SAM, which modifies logit statistics instead of function geometry.
result Functional-SAM and combined methods outperform AdamW and SAM in NLP tasks.
Soft diamond regularizers improve deep learning performance and sparsity.
problem Improving deep learning performance and sparsity of trained weights.
method New soft diamond synaptic weight priors based on thick-tailed symmetric alpha stable probability curves.
result Soft diamond regularizers outperform state-of-the-art methods in deep learning tasks.
A new memory-efficient sign language translation model reduces weight usage.
problem Memory constraints in real-time sign language translation.
method Variational Bayesian sequence-to-sequence network with Gaussian posterior and Indian Buffet Process prior.
result The proposed model achieves substantial weight compression without compromising performance.
ATD measures language distance using neural models, recovering linguistic groupings.
problem Lack of a unified quantitative measure for cross-linguistic distance.
method Pretrained multilingual language models, attention mechanisms, optimal transport.
result ATD quantifies representational distance between languages, recovering linguistic groupings.
Efficient echo state network with explicit memory performs well on benchmark tasks.
problem Training differentiable neural computers is difficult and time-consuming.
method Echo state network with an explicit memory.
result Echo state network can recognize all regular languages, including those contractive networks cannot.
This paper addresses the problem of inferring a regular expression from a given set of strings that resembles, as closely as possible, the regular expression that a human expert would have written to identify the language. This is motivated by our goal of automating the task of postmasters of an email service who use r…
Paper improves neural network models for MOOC student course prediction.
problem Improving neural network-based predictive methods for MOOC student course trajectory modeling.
method Enhanced LSTM networks and Transformer architecture applied to MOOC student data.
result Improved accuracy in predicting MOOC student course trajectories.
We survey - by means of 20 examples - the concept of varifold, as generalised submanifold, with emphasis on regularity of integral varifolds with mean curvature, while keeping prerequisites to a minimum. Integral varifolds are the natural language for studying the variational theory of the area integrand if one conside…
Aligns text representations over time for better performance in sequential tasks.
problem Language evolution causes data drift in sequential tasks.
method Sequentially aligns learned representations to combat data drift.
result Sequential alignment outperforms strong baselines on various tasks.
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA) method, but only on the subsequence of training examples where a classification error is made. We derive a bound on the number of mistakes…
ETC improves Transformer models for long and structured inputs.
problem Scaling input length and encoding structured inputs in Transformers.
method Introduces global-local attention, relative position encodings, and CPC pre-training.
result Achieves state-of-the-art results on four natural language datasets.
RotationOut rotates input vectors to regularize neural networks.
problem Reduction of co-adaptation in neural networks.
method Randomly rotates input vectors of the input layer.
result RotationOut reduces co-adaptation better than Dropout.