arXiv AI By Philip Quirke

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

Read the original on arXiv AI →

arXiv:2606. 08292v1 Announce Type: new Abstract: In mechanistic interpretability, attention heads are commonly elevated to role claims (e.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

The paper investigates the reliability of attention‑head ablation as a causal inference tool in language models. Using GPT‑2 small, the authors find that a natural post‑projection zeroing method is almost uncorrelated with a corrected pre‑projection ablation and yields a completely different set of top‑5 important heads. They also show that binary accuracy can mask effects near performance floors or ceilings, whereas gold‑token log‑probability provides a graded signal. By employing a discovery/held‑out split and 1,000 matched random‑head and layer‑matched‑head controls, the corrected per‑head effect ranking remains highly stable (Spearman ρ = 0.974) and the top‑5 heads significantly outperform both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is weak on GPT‑2, and replication on DistilGPT‑2 confirms the intervention‑semantic and matched‑control findings.

By Juli Huang
arXiv AI
Sep 12

How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism

Instruction tuning is often thought to give language models a universal ability to follow instructions, but this study shows otherwise. By probing nine tasks across three models, the authors find that general probes reveal selective, not uniform, deficits, cross‑task transfer is weak and skill‑similar, and causal ablation uncovers sparse, asymmetric dependencies. The results suggest instruction following is a coordinated use of diverse linguistic skills rather than a single shared mechanism.

By Elisabetta Rocchetti, Alfio Ferrara
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu