arXiv:2512. 20661v2 Announce Type: replace Abstract: Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, forcing models to over-focus on task-irrelevant tokens.
By Yawei Liu
arXiv:2609.28117v1 Announce Type: cross
Abstract: In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps....
By Pawe{\l} M\k{a}ka, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
arXiv:2604. 22583v2 Announce Type: replace Abstract: Multi-head attention enables Transformers to capture diverse representations, but all attention heads are typically activated for every input, regardless of task complexity.
By Bilal Faye, Abdoulaye Mbaye, Hanane Azzag, Mustapha Lebbah
arXiv:2502. 08363v3 Announce Type: replace-cross Abstract: We present Top-Theta (Top-$\theta$) Attention, a training-free method for sparsifying transformer attention during inference.
By Konstantin Berestizshevsky, Renzo Andri, Lukas Cavigelli
arXiv:2609.35868v1 Announce Type: new
Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations ca...
By Jinhao Zhang, Zeyu Liu, Zicheng Yan, Yunquan Zhang, Daning Cheng, Song Tang
arXiv:2512.20636v2 Announce Type: replace-cross
Abstract: Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppr...
By Dhananjay Saikumar, Blesson Varghese