A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Transformers can approximate any sequence-to-sequence function, surprising given their complexity.
problem Understanding the expressive power of Transformer models for sequence-to-sequence functions.
method Established that Transformers are universal approximators of continuous permutation equivariant sequence-to-sequence functions with compact support, and extended this to arbitrary functions using positional encodings.
result Transformers are universal approximators of arbitrary continuous sequence-to-sequence functions on a compact domain.
Single-head transformers with a single self-attention layer can approximate any sequence-to-sequence function and are efficient under certain conditions.
problem Statistical and computational limits of prompt tuning for transformer-based models.
method Investigation of single-head transformers with a single self-attention layer, proving universality and efficiency under SETH.
result Existence of almost-linear time prompt tuning inference algorithms under certain conditions.
Sparse Transformers can approximate dense Transformers with only O(n) connections.
problem Can sparse Transformers approximate arbitrary sequence-to-sequence functions?
method Proposed sufficient conditions for universal approximation and proved that sparse Transformers with O(n) connections can approximate dense models.
result Sparse Transformers with O(n) connections can approximate the same function class as dense models with n^2 connections.
We apply Cartan's method of equivalence to find a Bäcklund autotransformation for the tangent covering of the universal hierarchy equation. The transformation provides a recursion operator for symmetries of this equation.
Building on the universal covering group of the general linear group, we introduce the composite spinor bundle whose subbundles are Lorentz spin structures associated with different gravitational fields. General covariant transformations of this composite spinor bundle are canonically defined.
The paper proves a transformation theorem under a monotone property of almost Euclidean factors of geodesic balls.
problem The non-increasing property of numbers of almost Euclidean factors of geodesic balls.
method Proves a transformation theorem under a non-decreasing property compared to the non-increasing property.
result Shows that for a manifold with nonnegative Ricci curvature, if its universal cover is polar at infinity and the number of almost Euclidean factors is monotone, then its fundamental group is finitely generated and virtually abelian.
In the present paper, we study the finite type invariants of Gauss words. In the Polyak algebra techniques, we reduce the determination of the group structure to transformation of a matrix into its Smith normal form and we give the simplified form of a universal finite type invariant by means of the isomorphism of this…
Transformers can predict new tokens based on any number of context tokens, approximating continuous mappings with fixed resources.
problem Handling an arbitrarily large number of context tokens in transformers.
method Mathematical analysis of transformer's expressivity using Wasserstein distance and continuous mappings.
result Deep transformers are universal and can approximate continuous in-context mappings to arbitrary precision, uniformly over compact token domains.
The geometry of an admissible Bäcklund transformation for an exterior differential system is described by an admissible Cartan connection for a geometric structure on a tower with infinite--dimensional skeleton which is the universal prolongation of a ∣1∣--graded semi-simple Lie algebra.
Transformers can emulate various algorithms by prompting, proving universality.
problem How to emulate algorithms using fixed-weight Transformers.
method Two modes of in-context algorithm emulation: task-specific and prompt-programmable. Constructing prompts that encode algorithm parameters into token representations.
result Fixed-weight Transformers can emulate a broad class of algorithms via prompts.
We construct a certain cross product of two copies of the braided dual H~ of a quasitriangular Hopf algebra H, which we call the elliptic double EH, and which we use to construct representations of the punctured elliptic braid group extending the well-known representations of the planar braid group attache…
In this paper, we develop a theory about the relationship between G-invariant/equivariant functions and deep neural networks for finite group G. Especially, for a given G-invariant/equivariant function, we construct its universal approximator by deep neural network whose layers equip G-actions and each affine t…