arXiv Machine Learning By Bilal Faye, Abdoulaye Mbaye, Hanane Azzag, Mustapha Lebbah

Adaptive Head Budgeting for Efficient Multi-Head Attention

Read the original on arXiv Machine Learning →

arXiv:2604. 22583v2 Announce Type: replace Abstract: Multi-head attention enables Transformers to capture diverse representations, but all attention heads are typically activated for every input, regardless of task complexity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 28

Importance Scoring of Transformer Attention Heads in Learning Tabular Data

The paper introduces an importance‑scoring metric for multi‑head transformer attention heads applied to tabular data, a domain where transformers have been less studied. Experiments on 40 diverse tabular datasets show that removing heads with the lowest importance scores has minimal impact on performance, while removing the most important head first causes the largest drop. The study finds that important heads are distributed across layers and vary significantly across different tabular schemas, suggesting that the proposed score can help reduce redundancy and improve transformer efficiency.

By Ahmad Jad Allah, Kazi F. Akhter, Md. Kamrozzaman Bhuiyan, Manar D. Samad
arXiv Machine Learning
Sep 1

LoGo: Token-Level Dynamic Local-Global Attention

arXiv:2608.29539v1 Announce Type: cross Abstract: As context lengths scale, attention increasingly becomes a primary computational bottleneck in large language models. Standard Transformers remain po...

By Yuqi Pan, Zheng Li, Bohao Tang, Zhen Qin, Guoqi Li