Towards Verifiable Transformers: Solver-Checkable Circuit Explanations
arXiv:2605. 24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations.
arXiv:2605. 24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations.
arXiv:2510. 25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits.
arXiv:2607. 07316v1 Announce Type: new Abstract: This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks.
arXiv:2607. 00089v1 Announce Type: new Abstract: Mechanistic interpretability has produced a rich inventory of component-level analyses that characterise what neural-network components encode and how they interact.
arXiv:2509. 23806v2 Announce Type: replace-cross Abstract: Concolic testing for neural networks alternates concrete execution with constraint solving to search for inputs that flip model decisions.
arXiv:2608.31067v1 Announce Type: new Abstract: Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and lengt...
arXiv:2602.05896v3 Announce Type: replace-cross Abstract: Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in...
arXiv:2605. 18079v2 Announce Type: replace Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice.
arXiv:2608. 05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set.
The paper examines a method called Program-of-Layers (PoLar) that allows transformer layers to be dynamically routed rather than processed in a fixed sequence, mirroring the brain’s thalamic routing. Reproductions across five models confirm that skipping, repeating, and combining layer blocks improve performance, with shorter programs for easier inputs and more repeats for harder ones. However, the study could not replicate the claimed advantage of a learned single‑shot router, noting that its top prediction defaults to the standard pass while the top‑k predictions still yield accuracy gains. The authors also analyze the robustness of correction programs, finding them brittle to single edits, and release their code publicly.
arXiv:2607. 02964v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does.
arXiv:2609.15533v1 Announce Type: cross Abstract: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inheren...