arXiv Computation and Language By Richard Zhe Wang

The Communication Map of a Transformer

Read the original on arXiv Computation and Language →

The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 10

Through the Looking Glass: Directly Reading and Writing Transformers

The paper investigates how many transformer components influence a token prediction by measuring the absolute contribution of each unit and channel to the logit. It finds that thousands of components contribute to a single prediction, yet a small subset—often just dozens—carries the majority of the predictive mass. Across models ranging from 124 M to 7 B parameters, the proportion of the model involved in a prediction remains around one to three percent, independent of size, and the study demonstrates that specific components can be directly read and written to modify model behavior without additional training.

By Mark Oskin
arXiv AI
6d ago

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.

By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv AI
Aug 20

First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers

The paper introduces Language Identity Head Ablation (LIHA), a causal method that zeroes individual attention heads in transformer models to measure language switch rates across multilingual prompts. Applying LIHA to GPT‑2 reveals a small set of first‑token broadcaster heads—most notably L6H1—that persistently attend to the initial prompt token and propagate language signals throughout generation, with compensatory head recruitment occurring hierarchically in higher layers. A controlled comparison between Qwen2.5‑1.5B‑Base and Qwen2.5‑1.5B‑Instruct shows that instruction tuning concentrates language‑identity influence in early layers, while experiments with Chinese and Russian confirm script‑specific first‑token broadcasting at layer 0.

By Arjun Pillai, Christian Hoang, Anjelo Jann Laroza