arXiv Machine Learning By Athanasios Zeris

Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram

Read the original on arXiv Machine Learning →

arXiv:2606. 06573v1 Announce Type: cross Abstract: We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energetically dominant modes from turbulent flow ensembles.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 31

Critical attention scaling in long-context transformers

arXiv:2510. 05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length $n$ increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse.

By Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet