arXiv:2606. 16920v1 Announce Type: cross Abstract: Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task.
By Frank Zhengqing Wu, Francesco Tonin, Volkan Cevher
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.
By Vijay Erramilli
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, demonstrating that accurate average estimates can still lead to poor decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B show that observers trained on action loss tend to select lower‑loss actions, while traditional metrics like AUROC can rank monitors differently from deployment loss, highlighting the need for task‑specific evaluation.
arXiv:2601. 09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models.
By Jiali Cheng, Ziheng Chen, Chirag Agarwal, Hadi Amiri
arXiv:2607. 18921v1 Announce Type: cross Abstract: Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior.
By Yang Sheng, Jie Fu
arXiv:2609.36813v1 Announce Type: new
Abstract: Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. How...
By Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo