Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

164327491654 · Jun 202019922001200920172026
48 results for linear representations

The paper explores linear representations in language models using counterfactuals.

problem Understanding linear representations and geometric concepts in large language models.
method Formalized linear representation in output and input spaces, identified causal inner product.
result Unified understanding of linear representations and their connection to interpretation and control.

This work explains how linear representations in large language models arise from training objectives and gradient descent.

problem Understanding the origins of linear representations in large language models.
method A latent variable model to abstract and formalize concept dynamics, combined with analysis of the softmax cross-entropy objective and gradient descent.
result Linear representations emerge when learning from data matching the latent variable model, and this simple structure suffices to yield linear representations.

This paper explores the complexity of learning representations in contextual linear bandits.

problem Understanding the complexity of representation learning in contextual linear bandits.
method Systematic approach to representation learning in contextual linear bandits, focusing on instance-dependent perspective.
result Representation learning is fundamentally more complex than linear bandits, with some cases being arbitrarily harder.

Equivariant neural networks use symmetry to interpret complex data.

problem Interpreting and understanding the behavior of equivariant neural networks.
method Decompose layers into simple representations and analyze nonlinear activation functions.
result Equivariant neural networks can be interpreted using a filtration generalizing Fourier series.

Linear disentangled representations improve unsupervised action estimation.

problem Learning linear disentangled representations for unsupervised action estimation.
method Developed a method to induce irreducible representations in VAE models without labeled action sequences.
result Linear disentangled representations are a desirable property for unsupervised action estimation.

We consider deep feedforward neural networks with rectified linear units from a signal processing perspective. In this view, such representations mark the transition from using a single (data-driven) linear representation to utilizing a large collection of affine linear representations tailored to particular regions of…

2019-03-29abs ↗pdf ↗

New algorithms minimize regret in multi-task and lifelong linear bandits with shared representation.

problem Minimizing regret in multi-task and lifelong linear bandits with shared representation.
method Novel algorithms using efficient estimator for low-rank linear feature extractor and novel analysis.
result Achieved regret bounds matching minimax lower bound up to logarithmic factors.

This paper investigates how data augmentation improves linear separation of manifold data.

problem Understanding how data augmentation enhances linear separation of manifold data.
method Investigates the conditions under which self-supervised representations can linearly separate multi-manifold data.
result Self-supervised learning can linearly separate manifolds with a smaller distance than unsupervised learning.

Study of low dimensional representations of mapping class groups of surfaces, focusing on genus ≥ 7.

problem Classifying (2g+1)(2g+1)-dimensional complex linear representations of mapping class groups.
method Using twisted 1-cohomology groups and Morita's computation, a complete classification is given for g7g \geq 7.
result No irreducible linear representations of dimension 2g+12g+1 for g7g \geq 7.

Extends linear representation hypothesis to categorical and hierarchical concepts in LLMs.

problem Representing concepts without natural contrasts in large language models.
method Formalizes linear representation hypothesis for categorical and hierarchical concepts, proving relationships between concept hierarchy and representation geometry.
result Validated theoretical results on large language models, estimating representations for 900+ concepts.

We demonstrate that the notions of derivative representation of a Lie algebra on a vector bundle, of semi-linear representations of a Lie group on a vector bundle, and related concepts, may be understood in terms of representations of Lie algebroids and Lie groupoids, and we indicate how these notions extend to derivat…

2002-09-25abs ↗pdf ↗

It is a key to construct a similarity graph in graph-oriented subspace learning and clustering. In a similarity graph, each vertex denotes a data point and the edge weight represents the similarity between two points. There are two popular schemes to construct a similarity graph, i.e., pairwise distance based scheme an…

2013-04-24abs ↗pdf ↗

Paper proposes a method to learn linear regression models using multiple pre-trained models.

problem Learning a linear regression model with limited target data.
method Representation transfer learning method using multiple pre-trained models.
result The method achieves better sample complexity compared to baseline methods.

Study tight offline learning bounds for linear MDPs using variance information.

problem Understanding statistical limits with linear function representations in offline reinforcement learning.
method Variance-aware pessimistic value iteration (VAPVI) that reweights Bellman residuals based on estimated variances.
result Improved offline learning bounds expressed in terms of system quantities.

The paper shows that certain learned representations are identifiable in function space.

problem Identifiability of learned representations in deep neural networks.
method Using recent advances in nonlinear ICA, the paper shows that a large family of discriminative models are identifiable in function space, up to a linear indeterminacy.
result Many models for representation learning are identifiable in function space, including text, images, and audio.

Yu. I. Merzljakov developed a method of splittable coordinates which helps to verify the linearity of some groups, he established some fundamental results using this method. In this paper we use the method of splittable coordinates and find some sufficient condition under which the semi--direct product of two linear gr…

2005-06-07abs ↗pdf ↗

Meta-learning, or learning-to-learn, seeks to design algorithms that can utilize previous experience to rapidly learn new skills or adapt to new environments. Representation learning -- a key tool for performing meta-learning -- learns a data representation that can transfer knowledge across multiple tasks, which is es…

2020-02-26abs ↗pdf ↗

We construct analogues of FI-modules where the role of the symmetric group is played by the general linear groups and the symplectic groups over finite rings and prove basic structural properties such as Noetherianity. Applications include a proof of the Lannes--Schwartz Artinian conjecture in the generic representatio…

2014-08-16abs ↗pdf ↗

Improved texture synthesis using wavelet-based statistics with rectifier non-linearity.

problem Improving texture synthesis quality using wavelet representations.
method Proposes a family of statistics based on non-linear wavelet representations with a generalized rectifier non-linearity.
result Significantly improves visual quality of texture synthesis compared to classical wavelet-based models.

BCRL learns a Bellman complete representation for offline RL policy evaluation.

problem Learning a Q-function efficiently from offline data.
method BCRL learns a linear Bellman complete representation directly from data, enabling efficient OPE.
result BCRL achieves competitive OPE error and outperforms FQE in certain scenarios.

Linear representations help embed manifolds into matrix spaces.

problem Embedding manifolds into matrix spaces with effective bounds.
method Defining linear representations of G\mathsf{G}-manifolds as maps into matrix spaces, encoding G\mathsf{G}-actions as matrix products.
result Explicit bounds for Mostow-Palais G\mathsf{G}-equivariant embeddings of G\mathsf{G}-manifolds into G\mathsf{G}-modules V\mathbb{V}, showing dimV<\dim \mathbb{V} < \infty for compact G\mathsf{G}.

Study Anosov representations of reducible suspensions of hyperbolic groups.

problem Characterize dynamical properties of reducible suspensions of Anosov representations.
method Analyzing linear representations of non-elementary hyperbolic groups, focusing on weak unipotent actions on subspaces.
result Characterize when reducible suspensions are discrete and faithful, quasi-isometrically embedded, and Anosov.

The paper tackles learning from similar but not identical linear representations, improving performance over single-task learning.

problem Understanding how to learn from tasks with similar but not exactly the same linear representations, especially when dealing with outlier tasks.
method Proposes adaptive and robust penalized empirical risk minimization and spectral methods.
result Both methods outperform single-task learning when representations are similar and perform at least as well otherwise, with minimax optimality demonstrated.

Graph convolutional networks adapt the architecture of convolutional neural networks to learn rich representations of data supported on arbitrary graphs by replacing the convolution operations of convolutional neural networks with graph-dependent linear operations. However, these graph-dependent linear operations are d…

2017-11-03abs ↗pdf ↗

Classifies representations up to dimension 3g-3 for surface mapping class groups.

problem Classifying representations of mapping class groups up to a certain dimension.
method Direct sum of a 2g or 2g+1 dimensional representation and a trivial one.
result Any representation up to dimension 3g-3 is a direct sum of a 2g or 2g+1 dimensional representation and a trivial one.

Extending braid group representations to singular braid monoids and groups.

problem Extending braid group representations to singular braid monoids and groups.
method Investigating the extension of representations from braid groups to singular braid monoids and groups, and computing defects.
result Constructing a linear representation of the singular braid group that is an extension of the Lawrence-Krammer-Bigelow representation and computing its defect.

New models learn stable latent clusters without side info.

problem Stability of non-linear ICA representations without side information.
method Deep generative models with latent clusterings, compared to standard VAEs and auxiliary labeled models.
result Deep generative models with latent clusterings are as stable as models with side information.

Study multi-task learning with low-rank representation in stochastic linear bandits.

problem Transfer learning across multiple linear bandit tasks with shared low-dimensional representation.
method Proposes a greedy policy with trace norm regularization to implicitly learn a low-rank representation without knowing the rank.
result Upper bound on multi-task regret of NdT(T+d)r\sqrt{NdT(T+d)r}, showing benefit over independent task solving.

Low dimensional representations of words allow accurate NLP models to be trained on limited annotated data. While most representations ignore words' local context, a natural way to induce context-dependent representations is to perform inference in a probabilistic latent-variable sequence model. Given the recent succes…

2015-02-13abs ↗pdf ↗

This work defines idealized SSL representations and improves existing methods.

problem Unclear characteristics of SSL representations leading to high downstream accuracies.
method Characterized ideal properties and derived necessary and sufficient conditions.
result Improved SSL methods and derived new objectives for contrastive and non-contrastive learning.

Paper tackles intervention extrapolation using identifiable representations.

problem Predicting effects of unseen interventions on outcomes.
method Combines identifiable representation learning with autoencoders to enforce linear invariance.
result Identifiable representations enable non-linear extrapolation of interventions.

Good predictors of ICU Mortality have the potential to identify high-risk patients earlier, improve ICU resource allocation, or create more accurate population-level risk models. Machine learning practitioners typically make choices about how to represent features in a particular model, but these choices are seldom eva…

2015-12-16abs ↗pdf ↗

This paper considers exponential utility indifference pricing for a multidimensional non-traded assets model, and provides two linear approximations for the utility indifference price. The key tool is a probabilistic representation for the utility indifference price by the solution of a functional differential equation…

2014-03-30abs ↗pdf ↗

NGSLL combines DNN accuracy with linear model interpretability.

problem Combining high accuracy of DNNs with interpretability of linear models.
method Neural generators of sparse local linear models (NGSLL) using DNNs to approximate non-linear functions.
result Effective in real-world datasets, achieving high predictive performance and interpretability.

Study of modular representations in homology of congruence subgroups.

problem Understanding modular representations in homology of congruence subgroups.
method Analysis of sequences of modular representations of symplectic and special linear groups over finite fields.
result Established periodic representation stability in the sense of Church--Farb.

New algorithm improves learning efficiency in multi-task contextual bandits.

problem Improving learning efficiency in multi-task contextual bandits.
method Alternating projected gradient descent (GD) and minimization estimator for low-rank feature matrix recovery.
result Proved regret bound for multi-task learning algorithm.

We present a novel neural network algorithm, the Tensor Switching (TS) network, which generalizes the Rectified Linear Unit (ReLU) nonlinearity to tensor-valued hidden units. The TS network copies its entire input vector to different locations in an expanded representation, with the location determined by its hidden un…

2016-10-31abs ↗pdf ↗

FLAP adapts policies quickly to new tasks using shared linear representations.

problem Adapting policies to new tasks efficiently and effectively.
method FLAP uses a shared linear representation and a separate adapter network for quick adaptation.
result FLAP achieves up to 8X faster adaptation and significantly better performance on out-of-distribution tasks.