This paper proposes a new approach to Transformers by integrating hierarchical associative memory with MetaFormers.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
MLP-Mixer achieves better performance through sparsity and wider architecture.
MixerFlow combines MLP-Mixer with normalizing flows for efficient image modeling.
Interpolated-MLPs control inductive bias for better performance in low-compute tasks.
Vision Transformers show different internal representations compared to CNNs.
Study shows scaling up models doesn't always improve downstream tasks.
Bayesian sparsification reduces deep neural network complexity.
New method diagnoses criticality in deep neural networks, improving performance.
Transformers exhibit sparse activation maps, reducing computational load and improving robustness.