arXiv AI

First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers

The paper introduces Language Identity Head Ablation (LIHA), a causal method that zeroes individual attention heads in transformer models to measure language switch rates across multilingual prompts. Applying LIHA to GPT‑2 reveals a small set of first‑token broadcaster heads—most notably L6H1—that persistently attend to the initial prompt token and propagate language signals throughout generation, with compensatory head recruitment occurring hierarchically in higher layers. A controlled comparison between Qwen2.5‑1.5B‑Base and Qwen2.5‑1.5B‑Instruct shows that instruction tuning concentrates language‑identity influence in early layers, while experiments with Chinese and Russian confirm script‑specific first‑token broadcasting at layer 0.

arXiv AI
Sep 28

FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

The paper introduces FTB Graph, a method for mapping the causal circuitry that determines the first-token language identity in multilingual language models. Using Edge Attribution Patching and exact activation patching across six architectures (GPT‑2, BLOOM‑560M, Pythia‑1B/2.8B, Qwen2.5‑1.5B Base/Instruct), the authors extract directed acyclic graphs that reveal deep or mid‑to‑deep broadcasting hubs, with notable differences among models. The study finds that first‑token routing is largely established during pretraining and largely preserved by instruction tuning, while linear gradient approximations can diverge from causal interventions, underscoring the need for exact‑patching verification.

By Arjun Pillai, Christian Hoang, Anjelo Laroza
arXiv Computation and Language
Aug 25

The Communication Map of a Transformer

The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.

By Richard Zhe Wang
arXiv Computation and Language
Sep 25

Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.

By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv Computation and Language
Aug 25

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

arXiv:2608.12149v2 Announce Type: replace Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies:...

By Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo
arXiv AI
Sep 1

Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

The study investigates how post‑training of large autoregressive language models (ARMs) into masked diffusion models (MDMs) affects their internal computation. Across two 7B ARM‑MDM families and four diagnostic tasks, the authors find that MDMs retain much of the ARM’s high‑attribution pathways on prefix‑dominant tasks, but reorganize computation toward earlier layers on globally constrained tasks. Component‑level probes reveal that ARMs depend on sharply specialized components, whereas MDMs show weaker specialization and more diffuse output‑space alignment.

By Injin Kong, Hyoungjoon Lee, Yohan Jo
Hugging Face Trending Papers
Sep 3

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

arXiv Computation and Language
Sep 4

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

By Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
arXiv Machine Learning
5d ago

When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

The paper investigates the reliability of attention‑head ablation as a causal inference tool in language models. Using GPT‑2 small, the authors find that a natural post‑projection zeroing method is almost uncorrelated with a corrected pre‑projection ablation and yields a completely different set of top‑5 important heads. They also show that binary accuracy can mask effects near performance floors or ceilings, whereas gold‑token log‑probability provides a graded signal. By employing a discovery/held‑out split and 1,000 matched random‑head and layer‑matched‑head controls, the corrected per‑head effect ranking remains highly stable (Spearman ρ = 0.974) and the top‑5 heads significantly outperform both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is weak on GPT‑2, and replication on DistilGPT‑2 confirms the intervention‑semantic and matched‑control findings.

By Juli Huang