arXiv Machine Learning
Jun 8

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

arXiv:2604. 10098v2 Announce Type: replace Abstract: As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains.

By Zunhai Su, Hengyuan Zhang, Wei Wu, Yifan Zhang, Yaxiu Liu, He Xiao, Qingyao Yang, Yuxuan Sun, Rui Yang, Chao Zhang, Jing Xiong, Hui Shen, Keyu Fan, Weihao Ye, Chaofan Tao, Taiqiang Wu, Zhongwei Wan, Tiantian Zhang, Bowen Yan, Zhen Li, Yiming Zhang, Congkai Xie, Yulei Qian, Yuchen Xie, Yik-Chung Wu, Hongxia Yang, Ngai Wong
arXiv Computation and Language
Sep 7

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

The paper introduces an influence score that measures how much each attention head contributes to classification decisions in Transformer models, specifically for prompt injection detection. The score blends directional effects on logits with structural impact within the residual stream, allowing analysis at head, layer, and network scales. When applied to a DeBERTa model, the framework uncovers different decision patterns for correct versus incorrect predictions, offering a balanced approach between detailed circuit analysis and global output methods.

By Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi
arXiv Machine Learning
Jul 14

Gradient-Skipping Relevance Propagation for Efficient Explainability of Vision Transformers

arXiv:2607. 10365v1 Announce Type: cross Abstract: Vision Transformers (ViTs) are difficult to interpret because current methods of relevance propagation and attention flow do not fully consider some key architectural features, such as the uneven importance of attention heads and residual connections.

By Christopher Buratti, Michele Marchetti, Federica Parlapiano, Davide Traini, Domenico Ursino, Luca Virgili