Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

0111 · Jun 202019922001200920172026
4 results for Polysemanticity

New SAE algorithm proves feature recovery for LLMs with theoretical guarantees.

problem Achieving interpretable features in large language models (LLMs).
method Proposed a statistical framework and bias adaptation technique for sparse autoencoders (SAEs).
result Proved correct recovery of all monosemantic features under specific data sampling.

Describes explaining neurons in deep representations using compositional logical concepts.

problem Interpreting neuron behavior in deep neural networks.
method Identifying compositional logical concepts that closely approximate neuron behavior.
result Compositional explanations provide insights into model performance and allow for adversarial example creation.

AEN-SAEs address feature starvation in sparse autoencoders by stabilizing the geometric alignment of sparse coding.

problem Feature starvation in sparse autoencoders, leading to unstable and misaligned representations.
method Adaptive Elastic Net SAEs (AEN-SAEs) combine 2\ell_2 and 1\ell_1 terms to stabilize the sparse coding map and control feature interactions.
result AEN-SAEs mitigate feature starvation without heuristic resampling, maintaining competitive reconstruction abilities.