arXiv AI

When Attribution Patching Lies: Diagnosis and a Second-Order Correction

arXiv:2606. 09899v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior.

arXiv Machine Learning
Aug 5

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.

By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
arXiv AI
6d ago

FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

The paper introduces FTB Graph, a method for mapping the causal circuitry that determines the first-token language identity in multilingual language models. Using Edge Attribution Patching and exact activation patching across six architectures (GPT‑2, BLOOM‑560M, Pythia‑1B/2.8B, Qwen2.5‑1.5B Base/Instruct), the authors extract directed acyclic graphs that reveal deep or mid‑to‑deep broadcasting hubs, with notable differences among models. The study finds that first‑token routing is largely established during pretraining and largely preserved by instruction tuning, while linear gradient approximations can diverge from causal interventions, underscoring the need for exact‑patching verification.

By Arjun Pillai, Christian Hoang, Anjelo Laroza
arXiv Machine Learning
1d ago

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.

By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
arXiv Machine Learning
Sep 10

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

The paper introduces OSOL, a method for mitigating higher‑order interference in multi‑domain reinforcement learning. OSOL selects a focus domain each iteration, uses token‑level footprints from the previous checkpoint to rank rebound risk, and applies an adaptively scaled correction to the GRPO update. Experiments on Qwen3‑30B‑A3B show a 5.7% improvement over the best baseline without higher‑order differentiation.

By Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Guojun Yin, Wei Lin, Ran He
arXiv Computation and Language
Sep 4

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

By Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang
arXiv Machine Learning
Jul 14

Memory Savings at What Cost? A Study of Alternatives to Backpropagation

arXiv:2506. 21833v2 Announce Type: replace Abstract: Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing.

By Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan
arXiv Machine Learning
Sep 23

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.

By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts