The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.
By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
arXiv:2609.37616v1 Announce Type: new
Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on th...
By Abhinav Rajeev Kumar (Lossfunk), Paras Chopra (Lossfunk)
arXiv:2606. 18307v1 Announce Type: cross Abstract: Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs).
By Zefan Wang, Lincheng Li, Tianyu Yu, Yuan Yao
arXiv:2601. 21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive.
By Jianhui Chen, Yuzhang Luo, Liangming Pan
arXiv:2606. 29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training data shapes the high-level behavioral decisions a model learns to make.
By Reza Habibi, Darian Lee, Magy Seif El-Nasr
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose