A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs.
arXiv:2510. 25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits.
By Rabin Adhikari
The paper proposes a new architecture for masked language modeling that replaces the Transformer attention mechanism with a stack of low‑rank bottleneck autoencoders. Each autoencoder mixes information locally, across the full sequence, and across attention heads, compressing and reconstructing inputs without training‑dependent width. An iterative refinement process at masked positions pulls embeddings toward a weighted neighbor average and then projects them back onto the learned manifold, achieving comparable performance to BERT with roughly 1.9× fewer FLOPs and matching BERT on rare‑token performance through a frequency‑aware training schedule.
By Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
arXiv:2605. 18079v2 Announce Type: replace Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice.
By Moritz Br\"osamle, Stephan Eckstein
arXiv:2608. 15459v1 Announce Type: cross Abstract: Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose reasoning.
By Aditya Singh
The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.
By Eduardo Novaes Hering