arXiv AI By Marios Papamichalis, Regina Ruane

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Read the original on arXiv AI →

arXiv:2608. 14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.