arXiv:2601. 21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive.
By Jianhui Chen, Yuzhang Luo, Liangming Pan
arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.
By Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo, Chia-Tse Shao, Yingxiao Ye, Aobo Yang, Vivek Miglani, Nehal Bandi
arXiv:2608.24482v1 Announce Type: cross
Abstract: Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative appr...
By Hang Chen, Jiaying Zhu, Wenya Wang
arXiv:2608. 12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood.
By Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
arXiv:2604. 10098v2 Announce Type: replace Abstract: As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains.
By Zunhai Su, Hengyuan Zhang, Wei Wu, Yifan Zhang, Yaxiu Liu, He Xiao, Qingyao Yang, Yuxuan Sun, Rui Yang, Chao Zhang, Jing Xiong, Hui Shen, Keyu Fan, Weihao Ye, Chaofan Tao, Taiqiang Wu, Zhongwei Wan, Tiantian Zhang, Bowen Yan, Zhen Li, Yiming Zhang, Congkai Xie, Yulei Qian, Yuchen Xie, Yik-Chung Wu, Hongxia Yang, Ngai Wong
arXiv:2608. 16747v1 Announce Type: cross Abstract: Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors.
By Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks
arXiv:2606. 29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training data shapes the high-level behavioral decisions a model learns to make.
By Reza Habibi, Darian Lee, Magy Seif El-Nasr
The paper introduces a circuit‑grounded framework that links training‑dynamics‑based data valuation with mechanistic interpretability. It defines data quality along learnability, challenge, and alignment, identifies internal model circuits that control these utilities, and uses them as controllable interfaces for data generation. The authors present SAMS, a stage‑aware scheduling method that steers circuit‑guided data to match the model’s evolving optimization needs, achieving more diverse and effective data than prompt‑based baselines on multiple‑choice QA tasks.
By Nakyung Lee, Sangwoo Hong, Jungwoo Lee
The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2609.36448v1 Announce Type: new
Abstract: Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. Ho...
By Junze Deng, Daouda Sow, Sen Lin, Yingbin Liang
arXiv:2609.08216v1 Announce Type: new
Abstract: Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they break under distribution shift, and they...
By Arun Vignesh Malarkkan, Xinyuan Wang, Yanjie Fu
arXiv:2607. 03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic.
By Taras Kutsyk, Bartosz Zieli\'nski