Study finds phase transition in context-sensitive language model with short-range interactions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A two-step approach efficiently selects hyperparameters for FCMs.
Parallel recordings of neural spike counts have revealed the existence of context-dependent noise correlations in neural populations. Theories of population coding have also shown that such correlations can impact the information encoded by neural populations about external stimuli. Although studies have shown that the…
Lie's third theorem proven for Lie ∞-algebras.
Ghost points affect stability in finite difference schemes for diffusion equations.
Random Transformers behave like polynomial models in ICL with asymptotic growth.
We consider a planning problem where the dynamics and rewards of the environment depend on a hidden static parameter referred to as the context. The objective is to learn a strategy that maximizes the accumulated reward across all contexts. The new model, called Contextual Markov Decision Process (CMDP), can model a cu…
We tackle the problem of online reward maximisation over a large finite set of actions described by their contexts. We focus on the case when the number of actions is too big to sample all of them even once. However we assume that we have access to the similarities between actions' contexts and that the expected reward…
The paper proves how Transformers learn from context and generalize well.
Curves can bound only finitely many developable surfaces.
Transformers learn to solve modular arithmetic tasks by in-context learning and skill composition.
Continuum transformers learn operators in context via gradient descent.
Attention-only transformers learn from context via two stages of inference.
Neural attention (NA) has become a key component of sequence-to-sequence models that yield state-of-the-art performance in as hard tasks as abstractive document summarization (ADS) and video captioning (VC). NA mechanisms perform inference of context vectors; these constitute weighted sums of deterministic input sequen…
We consider a class of operator-induced norms, acting as finite-dimensional surrogates to the L2 norm, and study their approximation properties over Hilbert subspaces of L2 . The class includes, as a special case, the usual empirical norm encountered, for example, in the context of nonparametric regression in reproduci…
This work analyzes how transformers learn common linear regression tasks.
New model predicts drug effects across various cell types using causal imputation.
INFERS PDEs from data samples using learned context.
On a multi-assets Black-Scholes economy, we introduce a class of barrier options. In this model we apply a generalized reflection principle in a context of the finite reflection group acting on a Euclidean space to give a valuation formula and the semi-static hedge.
Researchers analyze neural process architectures and their representational capacities.
Automaton models are often seen as interpretable models. Interpretability itself is not well defined: it remains unclear what interpretability means without first explicitly specifying objectives or desired attributes. In this paper, we identify the key properties used to interpret automata and propose a modification o…
We propose a new statistical model for computational linguistics. Rather than trying to estimate directly the probability distribution of a random sentence of the language, we define a Markov chain on finite sets of sentences with many finite recurrent communicating classes and define our language model as the invarian…
We show in this note that the Sobolev Discrepancy introduced in Mroueh et al in the context of generative adversarial networks, is actually the weighted negative Sobolev norm , that is known to linearize the Wasserstein distance and plays a fundamental role in the dynamic formulation of…
Unified framework for studying softmax attention under large prompts.
New algorithm learns optimal decisions from imperfectly observed contexts.
Detects corruption in agentic models during execution.
Gradient descent biases linear models in next-token prediction towards data entropy.
This paper approaches the definition and properties of dynamic convex risk measures through the notion of a family of concave valuation operators satisfying certain simple and credible axioms. Exploring these in the simplest context of a finite time set and finite sample space, we find natural risk-transfer and time-co…
Study entropy bounds and finiteness for symmetric self-shrinkers.
In risk management, tail risks are of crucial importance. The assessment of risks should be carried out in accordance with the regulatory authority's requirement at high quantiles. In general, the underlying distribution function is unknown, the database is sparse, and therefore special tail models are used. Very often…
We prove extension theorems for several geometric properties such as asymptotic property C (APC), finite decomposition complexity (FDC), strict finite decomposition complexity (sFDC) which are weakenings of Gromov's finite asymptotic dimension (FAD). The context of all theorems is a finitely generated group with a …
Bayesian optimization tackles uncertainty in context variables.
We explicate a number of notions of algebraic laminations existing in the literature, particularly in the context of an exact sequence of hyperbolic groups. These laminations arise in different contexts: existence of Cannon-Thurston maps; closed geodesics exiting ends of manifolds; dual to …
We show that the Kuratowski imbedding of a Riemannian manifold in L^\infty, exploited in Gromov's proof of the systolic inequality for essential manifolds, admits an approximation by a (1+C)-bi-Lipschitz (onto its image), finite-dimensional imbedding for every C>0. Our key tool is the first variation formula thought of…
The paper compares inserting and stretching points for grid refinement near critical points.
Study on how attention in prompt-tuning affects large language models.
We discuss intrinsic aspects of Krupka's approach to finite-order variational sequences. We give intrinsic isomorphisms of the quotient subsheaves of the short finite-order variational sequence with sheaves of forms on jet spaces of suitable order, obtaining a new finite-order (short exact) variational sequence which i…
Paper proves no specific CMC hypersurfaces in hyperbolic space.
The paper proves stability in compact finite dimensional Alexandrov spaces using equivariant Gromov--Hausdorff convergence.
We generalize basic results relating the associated graded Lie algebra and the holonomy Lie algebra from finitely presented, commutator-relators groups to arbitrary finitely presented groups. In the process, we give an explicit formula for the cup-product in the cohomology of a finite 2-complex, and an algorithm for co…
Algorithm identifies best arm in piecewise stationary linear bandits with minimal samples.
Generalizes neural networks for infinite-dimensional mappings, including PDE solutions.
The paper extends Nambu-Poisson structures to infinite dimensions.
This work studies the class of algorithms for learning with side-information that emerge by extending generative models with embedded context-related variables. Using finite mixture models (FMM) as the prototypical Bayesian network, we show that maximum-likelihood estimation (MLE) of parameters through expectation-maxi…
In this article we study the K- and L-theory of groups acting on trees. We consider the problem in the context of the fibered isomorphism conjecture of Farrell and Jones. We show that in the class of residually finite groups it is enough to prove the conjecture for finitely presented groups with one end. Also, we deduc…
Transformers learn unseen tasks via prompts without fine-tuning.
The study explores ends in coarse homotopy of proper geodesic spaces.
We introduce a novel class of labeled directed acyclic graph (LDAG) models for finite sets of discrete variables. LDAGs generalize earlier proposals for allowing local structures in the conditional probability distribution of a node, such that unrestricted label sets determine which edges can be deleted from the underl…