arXiv:2509.12635v4 Announce Type: replace-cross
Abstract: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention sco...
By Yu Wang, Sheng Shen, R\'emi Munos, Hongyuan Zhan, Yuandong Tian
arXiv:2609.38109v1 Announce Type: cross
Abstract: The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of...
By Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge
arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
By Chungpa Lee, Jy-yong Sohn, Kangwook Lee
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
By Sajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter, Antonio Orvieto, Volkan Cevher
arXiv:2609.06712v2 Announce Type: replace
Abstract: Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequ...
By Zekun Zhang, Yixiang Cai, Yuxi Liu, Tengxu Sun, Tianle Liu, Zhoutong Wu, Haoyu Li, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan
arXiv:2607. 19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.
By Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li