Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2356 · Sep 202019922001200920182026
48 results for N-gram LMs

Improved neural language models trained with dynamic noise-contrastive estimation.

problem Training large-scale language models efficiently and avoiding overfitting.
method Dynamic Noise-Contrastive Estimation (DNCE) to train neural trans-dimensional random field language models.
result DNCE reduces training cost and improves model performance on large datasets.

LM-SNNs use lattice maps to classify and cluster images.

problem Image classification and clustering.
method Lattice map spiking neural networks with cooperative and competitive interactions, inhibition strategies, and biologically motivated learning rules.
result LM-SNNs effectively classify and cluster images using self-organized filters.

Improved neural TRF LMs for speech recognition with NCE and CNN integration.

problem Training inefficiency of neural TRF LMs on large training corpora.
method Reformulated TRFs, noise-contrastive estimation, CNN integration.
result Successful and efficient training on a 40x larger dataset with 1/3 training time and 4.7% WER reduction.

This paper investigates teacher hacking during language model distillation and proposes methods to mitigate it.

problem Teacher hacking during language model distillation, leading to suboptimal performance.
method A controlled experimental setup involving an oracle LM, teacher LM, and student LM, using fixed offline or online data generation techniques.
result Data diversity is the key factor in preventing teacher hacking during distillation.

New deep architectures inspired by numerical differential equations improve performance.

problem Designing effective deep neural networks.
method Interpreting deep networks as numerical discretizations of differential equations.
result LM-ResNet and LM-ResNeXt achieve higher accuracy with less parameters.

Let LM be the semigroup of non-degenerate based loops with a fixed initial/final frame in a Riemannian manifold M of dimension at least three. We compare the topology of LM to that of the loop space Omega FTM on the bundle of frames in the tangent bundle of M. We show that Omega FTM is the group completion of LM, and p…

2012-09-18abs ↗pdf ↗

We introduce a probabilistic approach to the LMS filter. By means of an efficient approximation, this approach provides an adaptable step-size LMS algorithm together with a measure of uncertainty about the estimation. In addition, the proposed approximation preserves the linear complexity of the standard LMS. Numerical…

2015-01-27abs ↗pdf ↗

We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably dominates its conventional counterpart in terms of mean square deviations. We es…

2010-12-22abs ↗pdf ↗

Let M be a closed, connected manifold, and LM its loop space. In this paper we describe closed string topology operations in h_*(LM), where h_* is a generalized homology theory that supports an orientation of M. We will show that these operations give h_*(LM) the structure of a unital, commutative Frobenius algebra wit…

2003-02-28abs ↗pdf ↗

This paper shows RL with KL penalties is equivalent to Bayesian inference for fine-tuning LMs.

problem Fine-tuning large language models to avoid undesirable features.
method Analyzed KL-regularized RL and showed it's equivalent to variational inference.
result KL-regularized RL avoids distribution collapse and is more insightful as Bayesian inference.

Neural TRFs improve speech recognition models with fewer parameters and faster inference.

problem Improving speech recognition models with fewer resources.
method Introducing neural TRFs that use nonlinear potentials with continuous features implemented by neural networks, combined with efficient inference techniques.
result Neural TRFs outperform discrete TRFs and LSTM LMs with fewer parameters and faster inference.

In this paper, we prove that for every Finsler nn-dimensional sphere (Sn,F)(S^{n},F) with reversibility $\lm$ and flag curvature KK satisfying $\left(\frac{\lm}{1+\lm}\right)^2<K\le 1$, either there exist infinitely many closed geodesics, or there exist at least two elliptic closed geodesics and each linearized Poincaré …

2015-04-01abs ↗pdf ↗

We propose a version of least-mean-square (LMS) algorithm for sparse system identification. Our algorithm called online linearized Bregman iteration (OLBI) is derived from minimizing the cumulative prediction error squared along with an l1-l2 norm regularizer. By systematically treating the non-differentiable regulariz…

2012-10-01abs ↗pdf ↗

New methods improve integration of external LMs with AED models.

problem Improving performance of AED models by integrating external LMs.
method Comparing and proposing novel methods to estimate implicit LM from AED models.
result Proposed methods outperform previous approaches.

LMs perform poorly in true few-shot learning without held-out examples.

problem Evaluating few-shot performance of language models without access to held-out examples.
method Evaluated two model selection criteria (cross-validation and minimum description length) for choosing LM prompts and hyperparameters in true few-shot learning.
result Selection criteria often prefer models that perform worse than random selection, suggesting overestimation of few-shot ability.

Using the Wodzicki residue, we build Wodzicki-Chern-Simons (WCS) classes in H2k1(LM)H^{2k-1}(LM) associated to the residue Chern character on the loop space LMLM of a Riemannian manifold M2k1M^{2k-1}. These WCS classes are associated to the L2L^2 connection and the Sobolev s=1s=1 connections on LM.LM. The WCS classes detect seve…

2014-07-09abs ↗pdf ↗

We present power low rank ensembles (PLRE), a flexible framework for n-gram language modeling where ensembles of low rank matrices and tensors are used to obtain smoothed probability estimates of words in context. Our method can be understood as a generalization of n-gram modeling to non-integer n, and includes standar…

2013-12-26abs ↗pdf ↗

Paper improves natural language understanding with less data using a new training method.

problem Limited data hinders performance of small models in natural language tasks.
method Generation-Distillation: uses large finetuned models to generate new training data and distill knowledge into smaller models.
result Achieves comparable performance to BERT with 300x fewer parameters and outperforms prior distillation methods.

Generative Adapter adapts LMs with a single forward pass, reducing inference overhead.

problem Efficient adaptation of large language models for new contexts.
method Generative Adapter directly maps new contexts to low-rank LM adapters via self-supervised learning.
result Significant reduction in inference overhead with no need for fine-tuning.

We study the existence of S1S^1-equivariant characteristic classes on certain natural infinite rank bundles over the loop space LMLM of a manifold MM. We discuss the different S1S^1-equivariant cohomology theories in the literature and clarify their relationships. We attempt to use S1S^1-equivariant Chern-Weil techniq…

2015-07-30abs ↗pdf ↗

Study shows challenges in converting RNNs to FSMs due to computational complexity.

problem Understanding the equivalence and distance between RNNs and FSMs.
method Computational proofs for equivalence and distance problems between RNNs and FSMs.
result Undecidability and hardness of approximation problems between RNNs and FSMs.

Given a smooth closed manifold M with a family {L_i} of closed submanifolds, we consider the free loop space LM and the spaces PM(L_i,L_j) of open strings (paths g:[0,1]->M with g(0) in L_i, and g(1) in L_j). We construct string topology operations resulting in an open-closed TQFT on the family (h_*(LM),h_*(PM(L_i,L_j)…

2006-06-20abs ↗pdf ↗

Chas and Sullivan have defined an intersection-type product on the homology of the free loop space LM of an oriented manifold M. In this paper we show how to extend this construction to a topological conformal field theory of degree d. In particular, we get operations on the homology of LM which are parameterized by th…

2007-11-30abs ↗pdf ↗

Paper shows how LSTM can remember long sequences by attending to persisted information.

problem LSTMs struggle with long sequences due to fading information and bias towards recent data.
method The paper introduces a mechanism that allows LSTMs to attend to information in memory based on how long it was persisted by the gating mechanism.
result The method improves LSTM's ability to process long sequences by retrieving information proportionally to its persistence in memory.

Let GPMG \to P \to M be a flat principal bundle over a closed and oriented manifold MM of dimension m=2dm=2d. We construct a map of Lie algebras $Ψ: \H_{2\ast} (L M) \to ø(\Mc)$, where $\H_{2\ast} (LM)$ is the even dimensional part of the equivariant homology of LMLM, the free loop space of MM, and $\Mc$ is the Maurer-C…

2006-02-06abs ↗pdf ↗

Let M be one of the projective spaces CP^n, HP^n for n>1 or the Cayley projective plane OP^2, and let LM denote the free loop space on M. Using Morse theory methods, we prove that the suspension spectrum of (LM)_+ is homotopy equivalent to the suspension spectrum of M_+ wedge a family of Thom spaces of explicit vector …

2005-11-03abs ↗pdf ↗

In LM, we proved a family version of the famous Witten rigidity theorems and several family vanishing theorems for elliptic genera. In this paper, we gerenalize our theorems LM in two directions. First we establish a family rigidity theorem for the Dirac operator on loop space twisted by general positive energy loop gr…

1999-11-05abs ↗pdf ↗

Let MM be a closed, oriented manifold of dimension dd. Let LMLM be the space of smooth loops in MM. Chas and Sullivan recently defined a product on the homology H(LM)H_*(LM) of degree d-d. They then investigated other structure that this product induces, including a Batalin -Vilkovisky structure, and a Lie algebra str…

2001-07-25abs ↗pdf ↗

Proposes CLRS-Text, a new benchmark for evaluating LM reasoning capabilities.

problem Lack of transferable benchmarks for evaluating reasoning capabilities of language models.
method Developed a textual version of the CLRS benchmark, generating diverse algorithmic tasks.
result Demonstrates a novel challenge for the LM reasoning community and validates prior work.