arXiv AI By Jianhui Chen, Yuzhang Luo, Liangming Pan

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

Read the original on arXiv AI →

arXiv:2601. 21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

The paper investigates how LLaMA 3.1‑8B models numerical sequence patterns, focusing on time‑series prediction. By designing a task that requires detecting structural cues—specifically first differences in a sequence—the authors show that the model performs well and internally computes and stores these differences. Probing and activation‑patching experiments reveal that LLaMA retrieves and applies the first‑difference via an induction‑like circuit, marking one of the first demonstrations of concept induction in large language models.

By Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang
arXiv AI
Aug 13

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

arXiv:2608. 12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood.

By Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen