A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
This work explains how linear representations in large language models arise from training objectives and gradient descent.
problem Understanding the origins of linear representations in large language models.
method A latent variable model to abstract and formalize concept dynamics, combined with analysis of the softmax cross-entropy objective and gradient descent.
result Linear representations emerge when learning from data matching the latent variable model, and this simple structure suffices to yield linear representations.
We consider deep feedforward neural networks with rectified linear units from a signal processing perspective. In this view, such representations mark the transition from using a single (data-driven) linear representation to utilizing a large collection of affine linear representations tailored to particular regions of…
Extends linear representation hypothesis to categorical and hierarchical concepts in LLMs.
problem Representing concepts without natural contrasts in large language models.
method Formalizes linear representation hypothesis for categorical and hierarchical concepts, proving relationships between concept hierarchy and representation geometry.
result Validated theoretical results on large language models, estimating representations for 900+ concepts.
We demonstrate that the notions of derivative representation of a Lie algebra on a vector bundle, of semi-linear representations of a Lie group on a vector bundle, and related concepts, may be understood in terms of representations of Lie algebroids and Lie groupoids, and we indicate how these notions extend to derivat…
It is a key to construct a similarity graph in graph-oriented subspace learning and clustering. In a similarity graph, each vertex denotes a data point and the edge weight represents the similarity between two points. There are two popular schemes to construct a similarity graph, i.e., pairwise distance based scheme an…
The paper shows that certain learned representations are identifiable in function space.
problem Identifiability of learned representations in deep neural networks.
method Using recent advances in nonlinear ICA, the paper shows that a large family of discriminative models are identifiable in function space, up to a linear indeterminacy.
result Many models for representation learning are identifiable in function space, including text, images, and audio.
Yu. I. Merzljakov developed a method of splittable coordinates which helps to verify the linearity of some groups, he established some fundamental results using this method. In this paper we use the method of splittable coordinates and find some sufficient condition under which the semi--direct product of two linear gr…
Meta-learning, or learning-to-learn, seeks to design algorithms that can utilize previous experience to rapidly learn new skills or adapt to new environments. Representation learning -- a key tool for performing meta-learning -- learns a data representation that can transfer knowledge across multiple tasks, which is es…
We construct analogues of FI-modules where the role of the symmetric group is played by the general linear groups and the symplectic groups over finite rings and prove basic structural properties such as Noetherianity. Applications include a proof of the Lannes--Schwartz Artinian conjecture in the generic representatio…
We propose a family of new representations of the braid groups on surfaces that extend linear representations of the braid groups on a disc such as the Burau representation and the Lawrence-Krammer-Bigelow representation.
Contrastive learning is an approach to representation learning that utilizes naturally occurring similar and dissimilar pairs of data points to find useful embeddings of data. In the context of document classification under topic modeling assumptions, we prove that contrastive learning is capable of recovering a repres…
The paper tackles learning from similar but not identical linear representations, improving performance over single-task learning.
problem Understanding how to learn from tasks with similar but not exactly the same linear representations, especially when dealing with outlier tasks.
method Proposes adaptive and robust penalized empirical risk minimization and spectral methods.
result Both methods outperform single-task learning when representations are similar and perform at least as well otherwise, with minimax optimality demonstrated.
Graph convolutional networks adapt the architecture of convolutional neural networks to learn rich representations of data supported on arbitrary graphs by replacing the convolution operations of convolutional neural networks with graph-dependent linear operations. However, these graph-dependent linear operations are d…
Extending braid group representations to singular braid monoids and groups.
problem Extending braid group representations to singular braid monoids and groups.
method Investigating the extension of representations from braid groups to singular braid monoids and groups, and computing defects.
result Constructing a linear representation of the singular braid group that is an extension of the Lawrence-Krammer-Bigelow representation and computing its defect.
Low dimensional representations of words allow accurate NLP models to be trained on limited annotated data. While most representations ignore words' local context, a natural way to induce context-dependent representations is to perform inference in a probabilistic latent-variable sequence model. Given the recent succes…
Good predictors of ICU Mortality have the potential to identify high-risk patients earlier, improve ICU resource allocation, or create more accurate population-level risk models. Machine learning practitioners typically make choices about how to represent features in a particular model, but these choices are seldom eva…
This paper considers exponential utility indifference pricing for a multidimensional non-traded assets model, and provides two linear approximations for the utility indifference price. The key tool is a probabilistic representation for the utility indifference price by the solution of a functional differential equation…
In the two parts of this paper we solve a problem of De Rham, proving that Reidemeister torsion invariants determine topological equivalence of linear G-representations, for G a finite cyclic group. Methods in controlled K-theory and surgery theory are developed to establish, and effectively calculate, a necessary and …
We present a novel neural network algorithm, the Tensor Switching (TS) network, which generalizes the Rectified Linear Unit (ReLU) nonlinearity to tensor-valued hidden units. The TS network copies its entire input vector to different locations in an expanded representation, with the location determined by its hidden un…