The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
By Po-Kai Chen, Aske Plaat, Niki van Stein
Nexus introduces a depth‑adaptive KV‑cache splicing and retrieval‑decoupled tool routing mechanism for agentic large language models that reduces the time‑to‑first‑token (TTFT) by decoupling tool routing from the expensive schema re‑encoding step. It uses an INT8 semantic lookaside buffer to select tools via retrieval and generates arguments from a compressed textual signature, maintaining about 89% routing accuracy even as the tool registry scales to 250 tools. Additionally, Nexus can splice compiled schema KV blocks into the live context, repairing the seam with a depth‑adaptive suffix redecode when rotary position embedding drift exceeds a threshold, ensuring output fidelity while achieving up to 1.7× TTFT speedup at moderate depth.
By Mustafa Arslan
arXiv:2605. 23393v2 Announce Type: replace-cross Abstract: Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2607. 25532v1 Announce Type: new Abstract: Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$).
By Athanasios Vlontzos, Giorgos Papanastasiou, Bernhard Kainz, Sotirios Tsaftaris
arXiv:2607. 24797v1 Announce Type: cross Abstract: In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and a fronto-parietal encoding route (impaired in pure agraphia), sharing a partial orthographic core.
By Diego Salda\~na Ulloa
arXiv:2606. 24948v1 Announce Type: new Abstract: Knowledge graph embedding (KGE) models predict single-hop links well but have no mechanism for zero-shot compositional queries: multi-hop questions whose relation chains never appeared during training.
By Randhir Kumar
arXiv:2607. 18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes?
By Jan Kirin
arXiv:2603. 13259v2 Announce Type: replace-cross Abstract: When a decoder-only transformer is forced to process matched correct and incorrect single-token continuations of a factual query, the two pathways through hidden-state space diverge in a specific way: displacement vectors from the query-only representation maintain approximately equal magnitude but rotate apart in direction.
By Javier Mar\'in
The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
By Liyun Zhang, Jiayi Guo
The paper investigates how many transformer components influence a token prediction by measuring the absolute contribution of each unit and channel to the logit. It finds that thousands of components contribute to a single prediction, yet a small subset—often just dozens—carries the majority of the predictive mass. Across models ranging from 124 M to 7 B parameters, the proportion of the model involved in a prediction remains around one to three percent, independent of size, and the study demonstrates that specific components can be directly read and written to modify model behavior without additional training.
By Mark Oskin
arXiv:2607. 09678v1 Announce Type: new Abstract: When LLM agents hand off information to one another, does the message format matter?
By Zayx Shawn
The paper audits whether routing entropy in Attention‑Residual transformer variants (Swin‑Tiny and DeiT‑Small) trained on CIFAR‑10/100 can signal prediction uncertainty beyond model confidence. Three tests examine the presence, consistency, and predictive power of routing signals, while a sensitivity audit measures how much injected effect the probes recover. Results show no significant improvement over confidence alone, with only modest recovery of injected signals and no consistent gains across seeds or metrics.
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen