arXiv AI By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

Read the original on arXiv AI →

The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

The paper investigates how training data attribution (TDA) can be used to influence large language models (LLMs). It compares two methods—reweighting and influence-guided response rewriting—on examples selected by influence functions. Rewriting, which replaces responses while keeping instructions fixed, yields stronger, more persistent, and bidirectional behavioral changes than reweighting, suggesting that the intervention value of influential samples is better realized through rewriting.

By Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan