arXiv Machine Learning

Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods

arXiv AI
Sep 10

Parity, Sensitivity, and Transformers

arXiv:2602.05896v3 Announce Type: replace-cross Abstract: Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in...

By Alexander Kozachinskiy, Tomasz Steifer, Przemys{\l}aw Wa{\l}\c{e}ga
arXiv AI
6d ago

Programs-of-Layers in LLMs through the Lens of Cortical Areas

The paper examines a method called Program-of-Layers (PoLar) that allows transformer layers to be dynamically routed rather than processed in a fixed sequence, mirroring the brain’s thalamic routing. Reproductions across five models confirm that skipping, repeating, and combining layer blocks improve performance, with shorter programs for easier inputs and more repeats for harder ones. However, the study could not replicate the claimed advantage of a learned single‑shot router, noting that its top prediction defaults to the standard pass while the top‑k predictions still yield accuracy gains. The authors also analyze the robustness of correction programs, finding them brittle to single edits, and release their code publicly.

By Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers