Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

1345 · Nov 201719922001200920172026
48 results for LM

This paper investigates teacher hacking during language model distillation and proposes methods to mitigate it.

problem Teacher hacking during language model distillation, leading to suboptimal performance.
method A controlled experimental setup involving an oracle LM, teacher LM, and student LM, using fixed offline or online data generation techniques.
result Data diversity is the key factor in preventing teacher hacking during distillation.

Let LM be the semigroup of non-degenerate based loops with a fixed initial/final frame in a Riemannian manifold M of dimension at least three. We compare the topology of LM to that of the loop space Omega FTM on the bundle of frames in the tangent bundle of M. We show that Omega FTM is the group completion of LM, and p…

2012-09-18abs ↗pdf ↗

We introduce a probabilistic approach to the LMS filter. By means of an efficient approximation, this approach provides an adaptable step-size LMS algorithm together with a measure of uncertainty about the estimation. In addition, the proposed approximation preserves the linear complexity of the standard LMS. Numerical…

2015-01-27abs ↗pdf ↗

We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably dominates its conventional counterpart in terms of mean square deviations. We es…

2010-12-22abs ↗pdf ↗

Trans-dimensional random field language models (TRF LMs) have recently been introduced, where sentences are modeled as a collection of random fields. The TRF approach has been shown to have the advantages of being computationally more efficient in inference than LSTM LMs with close performance and being able to flexibl…

2017-07-23abs ↗pdf ↗

Let M be a closed, connected manifold, and LM its loop space. In this paper we describe closed string topology operations in h_*(LM), where h_* is a generalized homology theory that supports an orientation of M. We will show that these operations give h_*(LM) the structure of a unital, commutative Frobenius algebra wit…

2003-02-28abs ↗pdf ↗

This paper shows RL with KL penalties is equivalent to Bayesian inference for fine-tuning LMs.

problem Fine-tuning large language models to avoid undesirable features.
method Analyzed KL-regularized RL and showed it's equivalent to variational inference.
result KL-regularized RL avoids distribution collapse and is more insightful as Bayesian inference.

In this paper, we prove that for every Finsler nn-dimensional sphere (Sn,F)(S^{n},F) with reversibility $\lm$ and flag curvature KK satisfying $\left(\frac{\lm}{1+\lm}\right)^2<K\le 1$, either there exist infinitely many closed geodesics, or there exist at least two elliptic closed geodesics and each linearized Poincaré …

2015-04-01abs ↗pdf ↗

We propose a version of least-mean-square (LMS) algorithm for sparse system identification. Our algorithm called online linearized Bregman iteration (OLBI) is derived from minimizing the cumulative prediction error squared along with an l1-l2 norm regularizer. By systematically treating the non-differentiable regulariz…

2012-10-01abs ↗pdf ↗

New methods improve integration of external LMs with AED models.

problem Improving performance of AED models by integrating external LMs.
method Comparing and proposing novel methods to estimate implicit LM from AED models.
result Proposed methods outperform previous approaches.

Using the Wodzicki residue, we build Wodzicki-Chern-Simons (WCS) classes in H2k1(LM)H^{2k-1}(LM) associated to the residue Chern character on the loop space LMLM of a Riemannian manifold M2k1M^{2k-1}. These WCS classes are associated to the L2L^2 connection and the Sobolev s=1s=1 connections on LM.LM. The WCS classes detect seve…

2014-07-09abs ↗pdf ↗

LMs perform poorly in true few-shot learning without held-out examples.

problem Evaluating few-shot performance of language models without access to held-out examples.
method Evaluated two model selection criteria (cross-validation and minimum description length) for choosing LM prompts and hyperparameters in true few-shot learning.
result Selection criteria often prefer models that perform worse than random selection, suggesting overestimation of few-shot ability.

Paper improves natural language understanding with less data using a new training method.

problem Limited data hinders performance of small models in natural language tasks.
method Generation-Distillation: uses large finetuned models to generate new training data and distill knowledge into smaller models.
result Achieves comparable performance to BERT with 300x fewer parameters and outperforms prior distillation methods.

Generative Adapter adapts LMs with a single forward pass, reducing inference overhead.

problem Efficient adaptation of large language models for new contexts.
method Generative Adapter directly maps new contexts to low-rank LM adapters via self-supervised learning.
result Significant reduction in inference overhead with no need for fine-tuning.

We study the existence of S1S^1-equivariant characteristic classes on certain natural infinite rank bundles over the loop space LMLM of a manifold MM. We discuss the different S1S^1-equivariant cohomology theories in the literature and clarify their relationships. We attempt to use S1S^1-equivariant Chern-Weil techniq…

2015-07-30abs ↗pdf ↗

Study shows challenges in converting RNNs to FSMs due to computational complexity.

problem Understanding the equivalence and distance between RNNs and FSMs.
method Computational proofs for equivalence and distance problems between RNNs and FSMs.
result Undecidability and hardness of approximation problems between RNNs and FSMs.

Given a smooth closed manifold M with a family {L_i} of closed submanifolds, we consider the free loop space LM and the spaces PM(L_i,L_j) of open strings (paths g:[0,1]->M with g(0) in L_i, and g(1) in L_j). We construct string topology operations resulting in an open-closed TQFT on the family (h_*(LM),h_*(PM(L_i,L_j)…

2006-06-20abs ↗pdf ↗

Chas and Sullivan have defined an intersection-type product on the homology of the free loop space LM of an oriented manifold M. In this paper we show how to extend this construction to a topological conformal field theory of degree d. In particular, we get operations on the homology of LM which are parameterized by th…

2007-11-30abs ↗pdf ↗

Let GPMG \to P \to M be a flat principal bundle over a closed and oriented manifold MM of dimension m=2dm=2d. We construct a map of Lie algebras $Ψ: \H_{2\ast} (L M) \to ø(\Mc)$, where $\H_{2\ast} (LM)$ is the even dimensional part of the equivariant homology of LMLM, the free loop space of MM, and $\Mc$ is the Maurer-C…

2006-02-06abs ↗pdf ↗

Let M be one of the projective spaces CP^n, HP^n for n>1 or the Cayley projective plane OP^2, and let LM denote the free loop space on M. Using Morse theory methods, we prove that the suspension spectrum of (LM)_+ is homotopy equivalent to the suspension spectrum of M_+ wedge a family of Thom spaces of explicit vector …

2005-11-03abs ↗pdf ↗

In LM, we proved a family version of the famous Witten rigidity theorems and several family vanishing theorems for elliptic genera. In this paper, we gerenalize our theorems LM in two directions. First we establish a family rigidity theorem for the Dirac operator on loop space twisted by general positive energy loop gr…

1999-11-05abs ↗pdf ↗

Let MM be a closed, oriented manifold of dimension dd. Let LMLM be the space of smooth loops in MM. Chas and Sullivan recently defined a product on the homology H(LM)H_*(LM) of degree d-d. They then investigated other structure that this product induces, including a Batalin -Vilkovisky structure, and a Lie algebra str…

2001-07-25abs ↗pdf ↗

Proposes CLRS-Text, a new benchmark for evaluating LM reasoning capabilities.

problem Lack of transferable benchmarks for evaluating reasoning capabilities of language models.
method Developed a textual version of the CLRS benchmark, generating diverse algorithmic tasks.
result Demonstrates a novel challenge for the LM reasoning community and validates prior work.

A new heuristic LM algorithm improves kk-segmentation accuracy with less computation.

problem Efficiently segmenting large video streams into meaningful piecewise-linear segments.
method Inspired by Lloyd's and Lloyd-Max algorithms, LM algorithm iteratively minimizes a cost function.
result LM algorithm achieves competitive accuracy with exact methods at a fraction of the computational cost.

We extend finite dimensional Chern-Simons theory to certain infinite dimensional principal bundles with connections, in particular to the frame bundle FLMLMFLM\to LM over the loop space of a Riemannian manifold MM. Chern-Simons forms are defined roughly as in finite dimensions with the invariant polynomials replaced by a…

2004-11-08abs ↗pdf ↗

Let MM be a compact oriented dd-dimensional smooth manifold and XX a topological space. Chas and Sullivan \cite{Chas-Sullivan:stringtop} have defined a structure of Batalin-Vilkovisky algebra on H(LM):=H+d(LM)\mathbb{H}_*(LM):=H_{*+d}(LM). Getzler \cite{Getzler:BVAlg} has defined a structure of Batalin-Vilkovisky algebra on the…

2009-08-13abs ↗pdf ↗

Study examines how decoding algorithms affect fairness in language generation models.

problem Impact of decoding algorithms on fairness in open-ended language generation.
method Systematic analysis of top-pp, top-kk, and temperature decoding algorithms.
result Decoding algorithms significantly impact fairness across demographic groups.

ProtTrans models predict protein features without evolutionary info.

problem Predicting protein features from amino acid sequences.
method Self-supervised deep learning on large protein datasets.
result ProtT5 embeddings outperform state-of-the-art for per-residue predictions.

This paper introduces a gradient analysis framework to improve language model performance by rewarding good examples and penalizing bad ones.

problem Improving language model output quality by penalizing bad examples.
method Gradient analysis of loss functions to reward good examples and penalize bad ones.
result ExMATE is superior to MLE and combining DPO with ExMATE enhances performance.

Identifies conjugate points in spherical harmonics solutions of quasi-geostrophic equations.

problem Locating conjugate points in spherical harmonics solutions.
method Utilizing structure constants and quasi-geostrophic equations on the sphere, identifying conjugate points.
result Existence and location of conjugate points along spherical harmonics solutions.

LM optimization outperforms other methods in deep learning tasks but at high computational cost.

problem Finding efficient optimization methods for deep learning models.
method Comparing first-order (CG, SGD, LM, L-BFGS) and higher-order optimization functions.
result Levemberg-Marquardt (LM) optimization significantly improves convergence but at a high computational cost.