GPT-f uses language models to find new proofs in formal math.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Self-supervised skip-tree training improves mathematical reasoning in language models.
Stan is a probabilistic programming language that has been increasingly used for real-world scalable projects. However, to make practical inference possible, the language sacrifices some of its usability by adopting a block syntax, which lacks compositionality and flexible user-defined functions. Moreover, the semantic…
Study shows LLMs can extrapolate rules from out-of-distribution prompts.
Extends linear representation hypothesis to categorical and hierarchical concepts in LLMs.
In this paper, we consider formal series associated with events, profiles derived from events, and statistical models that make predictions about events. We prove theorems about realizations for these formal series using the language and tools of Hopf algebras.
This work explains how linear representations in large language models arise from training objectives and gradient descent.
Language models help text classification tasks by predicting next words.
This work offers a broad perspective on probabilistic modeling and inference in light of recent advances in probabilistic programming, in which models are formally expressed in Turing-complete programming languages. We consider a typical workflow and how probabilistic programming languages can help to automate this wor…
When a bilingual student learns to solve word problems in math, we expect the student to be able to solve these problem in both languages the student is fluent in,even if the math lessons were only taught in one language. However, current representations in machine learning are language dependent. In this work, we pres…
The paper explores linear representations in language models using counterfactuals.
We build deep RL agents that execute declarative programs expressed in formal language. The agents learn to ground the terms in this language in their environment, and can generalize their behavior at test time to execute new programs that refer to objects that were not referenced during training. The agents develop di…
We prove a Darboux theorem for formal deformations of Hamiltonian operators of hydrodynamic type (Dubrovin-Novikov). Not all deformations are equivalent to the original operator: there is a moduli 2-stack of normal forms. The paper utilizes three main concepts: 1) dg Lie algebras concentrated in degrees [-1,\infty) suc…
VALC provides concept-level interpretations of FLMs, overcoming word-level limitations.
LLMs are compared to Markov chains for natural language processing.
Graph theory provides a language for studying the structure of relations, and it is often used to study interactions over time too. However, it poorly captures the both temporal and structural nature of interactions, that calls for a dedicated formalism. In this paper, we generalize graph concepts in order to cope with…
A framework uses free probability to analyze Transformer models.
We use the supergeometric formalism, more precisely, the so-called "big bracket" (for which brackets and anchors are encoded by functions on some graded symplectic manifold) to address the theory of Jacobi algebroids and bialgebroids (following mainly Iglesias-Marrero and Grabowski-Marmo as a guideline). This formalism…
We develop the multilingual topic model for unaligned text (MuTo), a probabilistic model of text that is designed to analyze corpora composed of documents in two languages. From these documents, MuTo uses stochastic EM to simultaneously discover both a matching between the languages and multilingual latent topics. We d…
In this note we review the basic mathematical ideas used in finance in the language of modern physics. We focus on discrete time formalism, derive path integral and Green's function formulas for pricing. We also discuss various risk mitigation methods.
Transformers can be hard to interpret due to complex optima.
We introduce a formal language IE that is a variant of the language PAL developed in [van Benthem 2011] by adding a belief operator and a common belief operator,specializing to stochastic analysis. A constant symbol in the language denotes a stochastic process so that we can represent several financial events as formul…
New algebraic formalism for differential calculus in Diolic algebras.
In order to satisfy safety conditions, an agent may be constrained from acting freely. A safe controller can be designed a priori if an environment is well understood, but not when learning is employed. In particular, reinforcement learned (RL) controllers require exploration, which can be hazardous in safety critical …
This note, in a rather expository manner, serves as a conceptional introduction to the certain underlying mathematical structures encoding the geometric quantization formalism and the construction of Witten's quantum invariants, which is in fact organized in the language topological quantum field theory.
Newtonian, Lagrangian, and Hamiltonian dynamical systems are well formalized mathematically. They give rise to geometric structures describing motion of a point in smooth manifolds. Riemannian metric is a different geometric structure formalizing concepts of length and angle. The interplay of Riemannian metric and its …
Hamiltonian Monte Carlo (HMC) is arguably the dominant statistical inference algorithm used in most popular "first-order differentiable" Probabilistic Programming Languages (PPLs). However, the fact that HMC uses derivative information causes complications when the target distribution is non-differentiable with respect…
Neural architecture search methods are able to find high performance deep learning architectures with minimal effort from an expert. However, current systems focus on specific use-cases (e.g. convolutional image classifiers and recurrent language models), making them unsuitable for general use-cases that an expert migh…
We present an original theorem in auction theory: it specifies general conditions under which the sum of the payments of all bidders is necessarily not identically zero, and more generally not constant. Moreover, it explicitly supplies a construction for a finite minimal set of possible bids on which such a sum is not …
Valid certifies LLMs' domain adherence, bounding out-of-domain behavior.
We develop a new Low-level, First-order Probabilistic Programming Language (LF-PPL) suited for models containing a mix of continuous, discrete, and/or piecewise-continuous variables. The key success of this language and its compilation scheme is in its ability to automatically distinguish parameters the density functio…
Introduces topological deep learning for neural network classification problems.
Introduces a new geometric framework for non-perturbative BV-theory.
The EMNLP 2018 workshop BlackboxNLP was dedicated to resources and techniques specifically developed for analyzing and understanding the inner-workings and representations acquired by neural models of language. Approaches included: systematic manipulation of input to neural networks and investigating the impact on thei…
Improved language models by incorporating statistical discriminators.
Advocates for user-friendly RL problem descriptions to improve usability and generalization.
I introduce a new geometrical approach to thermo--statistical mechanics. Here I highlight the main physical ideas, and how do they translate into geometrical language. I contrast the present approach with previous thermo--statistical--geometrical formalisms, (pseudo-)Riemannian [Weinhold 1975; Ruppeiner 1979] as well a…
LLMs can help explain credit risk models but not autonomously.
New methods adapt conformal prediction to unknown subpopulation shifts.
The paper proves impossibilities and positive results for universal machine translation.
This paper uses Gaussian mixtures to mimic interactions in large language models.
It is notoriously difficult to control the behavior of reinforcement learning agents. Agents often learn to exploit the environment or reward signal and need to be retrained multiple times. The multi-objective reinforcement learning (MORL) framework separates a reward function into several objectives. An ideal MORL age…
Study improves estimation of rare language model outputs.
We give a physical derivation of generalized Kahler geometry. Starting from a supersymmetric nonlinear sigma model, we rederive and explain the results of Gualtieri regarding the equivalence between generalized Kahler geometry and the bi-hermitean geometry of Gates-Hull-Rocek. When cast in the language of supersymmetri…
Triangulation filters spurious circuits in multilingual models.
Deep learning is currently the subject of intensive study. However, fundamental concepts such as representations are not formally defined -- researchers "know them when they see them" -- and there is no common language for describing and analyzing algorithms. This essay proposes an abstract framework that identifies th…
Motion planning and control are key problems in a collection of robotic applications including the design of autonomous agile vehicles and of minimalist manipulators. These problems can be accurately formalized within the language of affine connections and of geometric control theory. In this paper we overview recent r…
Tensor logic aims to unify AI types with scalable and transparent features.