arXiv Machine Learning By Ananth Eswar, Pratinav Seth, Utsav Avaiya, Vinay Kumar Sankarapu

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

Read the original on arXiv Machine Learning →

arXiv:2607. 05355v1 Announce Type: cross Abstract: Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.