Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
Read the original on arXiv Computation and Language →The paper introduces an influence score that measures how much each attention head contributes to classification decisions in Transformer models, specifically for prompt injection detection. The score blends directional effects on logits with structural impact within the residual stream, allowing analysis at head, layer, and network scales. When applied to a DeBERTa model, the framework uncovers different decision patterns for correct versus incorrect predictions, offering a balanced approach between detailed circuit analysis and global output methods.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.