Improved speech recognition model with better performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recurrent neural network (RNN) language models (LMs) and Long Short Term Memory (LSTM) LMs, a variant of RNN LMs, have been shown to outperform traditional N-gram LMs on speech recognition tasks. However, these models are computationally more expensive than N-gram LMs for decoding, and thus, challenging to integrate in…
This paper investigates teacher hacking during language model distillation and proposes methods to mitigate it.
Trans-dimensional random field language models (TRF LMs) where sentences are modeled as a collection of random fields, have shown close performance with LSTM LMs in speech recognition and are computationally more efficient in inference. However, the training efficiency of neural TRF LMs is not satisfactory, which limit…
In our work, we bridge deep neural network design with numerical differential equations. We show that many effective networks, such as ResNet, PolyNet, FractalNet and RevNet, can be interpreted as different numerical discretizations of differential equations. This finding brings us a brand new perspective on the design…
Let LM be the semigroup of non-degenerate based loops with a fixed initial/final frame in a Riemannian manifold M of dimension at least three. We compare the topology of LM to that of the loop space Omega FTM on the bundle of frames in the tangent bundle of M. We show that Omega FTM is the group completion of LM, and p…
We introduce a probabilistic approach to the LMS filter. By means of an efficient approximation, this approach provides an adaptable step-size LMS algorithm together with a measure of uncertainty about the estimation. In addition, the proposed approximation preserves the linear complexity of the standard LMS. Numerical…
A new whole-sentence language model - neural trans-dimensional random field language model (neural TRF LM), where sentences are modeled as a collection of random fields, and the potential function is defined by a neural network, has been introduced and successfully trained by noise-contrastive estimation (NCE). In this…
We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably dominates its conventional counterpart in terms of mean square deviations. We es…
Neural language modeling (LM) has led to significant improvements in several applications, including Automatic Speech Recognition. However, they typically require large amounts of training data, which is not available for many domains and languages. In this study, we propose a multilingual neural language model archite…
Trans-dimensional random field language models (TRF LMs) have recently been introduced, where sentences are modeled as a collection of random fields. The TRF approach has been shown to have the advantages of being computationally more efficient in inference than LSTM LMs with close performance and being able to flexibl…
Let M be a closed, connected manifold, and LM its loop space. In this paper we describe closed string topology operations in h_*(LM), where h_* is a generalized homology theory that supports an orientation of M. We will show that these operations give h_*(LM) the structure of a unital, commutative Frobenius algebra wit…
This paper shows RL with KL penalties is equivalent to Bayesian inference for fine-tuning LMs.
In this paper, we prove that for every Finsler -dimensional sphere with reversibility $\lm$ and flag curvature satisfying $\left(\frac{\lm}{1+\lm}\right)^2<K\le 1$, either there exist infinitely many closed geodesics, or there exist at least two elliptic closed geodesics and each linearized Poincaré …
We propose a version of least-mean-square (LMS) algorithm for sparse system identification. Our algorithm called online linearized Bregman iteration (OLBI) is derived from minimizing the cumulative prediction error squared along with an l1-l2 norm regularizer. By systematically treating the non-differentiable regulariz…
Paper proves a spinorial version of Aubin's estimate for the Yamabe problem.
Let M be a closed, oriented, n -manifold, and LM its free loop space. Chas and Sullivan defined a commutative algebra structure in the homology of LM, and a Lie algebra structure in its equivariant homology. These structures are known as the string topology loop product and string bracket, respectively. In this paper w…
New methods improve integration of external LMs with AED models.
The purpose of this paper is to indicate that the recently proposed Momentum fractional least mean squares (mFLMS) algorithm has some serious flaws in its design and analysis. Our apprehensions are based on the evidence we found in the derivation and analysis in the paper titled: \textquotedblleft \textit{Momentum frac…
The dominant language models (LMs) such as n-gram and neural network (NN) models represent sentence probabilities in terms of conditionals. In contrast, a new trans-dimensional random field (TRF) LM has been recently introduced to show superior performances, where the whole sentence is modeled as a random field. In thi…
Using the Wodzicki residue, we build Wodzicki-Chern-Simons (WCS) classes in associated to the residue Chern character on the loop space of a Riemannian manifold . These WCS classes are associated to the connection and the Sobolev connections on The WCS classes detect seve…
LMs perform poorly in true few-shot learning without held-out examples.
Language Models (LMs) are important components in several Natural Language Processing systems. Recurrent Neural Network LMs composed of LSTM units, especially those augmented with an external memory, have achieved state-of-the-art results. However, these models still struggle to process long sequences which are more li…
Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences. Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences. However, they a…
Paper improves natural language understanding with less data using a new training method.
Generative Adapter adapts LMs with a single forward pass, reducing inference overhead.
We study the existence of -equivariant characteristic classes on certain natural infinite rank bundles over the loop space of a manifold . We discuss the different -equivariant cohomology theories in the literature and clarify their relationships. We attempt to use -equivariant Chern-Weil techniq…
Study shows challenges in converting RNNs to FSMs due to computational complexity.
Neural language models (LMs) based on recurrent neural networks (RNN) are some of the most successful word and character-level LMs. Why do they work so well, in particular better than linear neural LMs? Possible explanations are that RNNs have an implicitly better regularization or that RNNs have a higher capacity for …
Given a smooth closed manifold M with a family {L_i} of closed submanifolds, we consider the free loop space LM and the spaces PM(L_i,L_j) of open strings (paths g:[0,1]->M with g(0) in L_i, and g(1) in L_j). We construct string topology operations resulting in an open-closed TQFT on the family (h_*(LM),h_*(PM(L_i,L_j)…
New framework assesses LM uncertainty without thresholding.
Chas and Sullivan have defined an intersection-type product on the homology of the free loop space LM of an oriented manifold M. In this paper we show how to extend this construction to a topological conformal field theory of degree d. In particular, we get operations on the homology of LM which are parameterized by th…
Let be a flat principal bundle over a closed and oriented manifold of dimension . We construct a map of Lie algebras $Ψ: \H_{2\ast} (L M) \to ø(\Mc)$, where $\H_{2\ast} (LM)$ is the even dimensional part of the equivariant homology of , the free loop space of , and $\Mc$ is the Maurer-C…
Let M be one of the projective spaces CP^n, HP^n for n>1 or the Cayley projective plane OP^2, and let LM denote the free loop space on M. Using Morse theory methods, we prove that the suspension spectrum of (LM)_+ is homotopy equivalent to the suspension spectrum of M_+ wedge a family of Thom spaces of explicit vector …
In LM, we proved a family version of the famous Witten rigidity theorems and several family vanishing theorems for elliptic genera. In this paper, we gerenalize our theorems LM in two directions. First we establish a family rigidity theorem for the Dirac operator on loop space twisted by general positive energy loop gr…
In this work we give a comprehensive overview of the time consistency property of dynamic risk and performance measures, focusing on a the discrete time setup. The two key operational concepts used throughout are the notion of the LM-measure and the notion of the update rule that, we believe, are the key tools for stud…
Let be a closed, oriented manifold of dimension . Let be the space of smooth loops in . Chas and Sullivan recently defined a product on the homology of degree . They then investigated other structure that this product induces, including a Batalin -Vilkovisky structure, and a Lie algebra str…
Proposes CLRS-Text, a new benchmark for evaluating LM reasoning capabilities.
A new heuristic LM algorithm improves -segmentation accuracy with less computation.
Linguistic calibration improves long-form text confidence.
LMs inevitably generate hallucinations, but can be made statistically negligible.
We extend finite dimensional Chern-Simons theory to certain infinite dimensional principal bundles with connections, in particular to the frame bundle over the loop space of a Riemannian manifold . Chern-Simons forms are defined roughly as in finite dimensions with the invariant polynomials replaced by a…
Let be a compact oriented -dimensional smooth manifold and a topological space. Chas and Sullivan \cite{Chas-Sullivan:stringtop} have defined a structure of Batalin-Vilkovisky algebra on . Getzler \cite{Getzler:BVAlg} has defined a structure of Batalin-Vilkovisky algebra on the…
Study examines how decoding algorithms affect fairness in language generation models.
ProtTrans models predict protein features without evolutionary info.
This paper introduces a gradient analysis framework to improve language model performance by rewarding good examples and penalizing bad ones.
Identifies conjugate points in spherical harmonics solutions of quasi-geostrophic equations.
LM optimization outperforms other methods in deep learning tasks but at high computational cost.