arXiv AI By Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan

Scaling Interpretable Transformers with Parity Bottleneck Layers

Read the original on arXiv AI →

arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.