arXiv:2606. 08292v2 Announce Type: replace Abstract: Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information, and restoring that activation repairs the damage.
By Philip Quirke
The paper investigates the reliability of attention‑head ablation as a causal inference tool in language models. Using GPT‑2 small, the authors find that a natural post‑projection zeroing method is almost uncorrelated with a corrected pre‑projection ablation and yields a completely different set of top‑5 important heads. They also show that binary accuracy can mask effects near performance floors or ceilings, whereas gold‑token log‑probability provides a graded signal. By employing a discovery/held‑out split and 1,000 matched random‑head and layer‑matched‑head controls, the corrected per‑head effect ranking remains highly stable (Spearman ρ = 0.974) and the top‑5 heads significantly outperform both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is weak on GPT‑2, and replication on DistilGPT‑2 confirms the intervention‑semantic and matched‑control findings.
By Juli Huang
Instruction tuning is often thought to give language models a universal ability to follow instructions, but this study shows otherwise. By probing nine tasks across three models, the authors find that general probes reveal selective, not uniform, deficits, cross‑task transfer is weak and skill‑similar, and causal ablation uncovers sparse, asymmetric dependencies. The results suggest instruction following is a coordinated use of diverse linguistic skills rather than a single shared mechanism.
By Elisabetta Rocchetti, Alfio Ferrara
arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.
By Yongzhong Xu
arXiv:2605. 24059v2 Announce Type: replace Abstract: We present a three-step recipe for identifying attention-head circuits in pretrained transformers.
By Yongzhong Xu
arXiv:2609.38205v1 Announce Type: new
Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the tran...
By Muhammad Usama, Dong Eui Chang
arXiv:2606. 08105v1 Announce Type: new Abstract: When attention concentrates on a single token, a sink, what is the model actually computing?
By Lukas Fesser, Mozes Jacobs, Thomas Fel, Andy Keller, Sham Kakade
arXiv:2606. 05976v2 Announce Type: replace Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources.
By Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin, Bao Li, Jung-Hsien Chiang
The paper investigates how the upper spectral tails of weight matrices in decoder‑only transformer language models influence reasoning behavior. By performing controlled interventions on the query–key product and comparing them to factor‑level surgeries, the authors find that edits targeting the spectral tail more strongly affect model performance across multiple checkpoints and reasoning benchmarks. The study also explores how inverse participation ratios predict accuracy transitions and shows that tail‑aware low‑rank adaptations converge faster than standard methods.
By Ibne Farabi Shihab, Sanjida Akhter, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma
arXiv:2609.39971v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.
By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv:2605.06524v3 Announce Type: replace
Abstract: Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online sett...
By Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal