arXiv:2509.12635v4 Announce Type: replace-cross
Abstract: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention sco...
By Yu Wang, Sheng Shen, R\'emi Munos, Hongyuan Zhan, Yuandong Tian
arXiv:2607. 07678v1 Announce Type: new Abstract: Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly.
By Xinyi Wu, Siyuan Liu, Ali Jadbabaie
arXiv:2607. 19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.
By Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li
arXiv:2608.05154v3 Announce Type: replace
Abstract: Multimodal rotary positional encodings apply temporal, height, and width phases to interleaved text, image, and video tokens. This creates two ambi...
By Donggen Li
arXiv:2609.11913v1 Announce Type: new
Abstract: Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. He...
By Daniel Henrik Nevermann, Claudius Gros
arXiv:2512. 14391v3 Announce Type: replace-cross Abstract: In-context learning is fundamental to modern Large Language Models (LLMs); however, prevailing architectures impose a rigid and fixed contextual structure by assigning linear or constant positional indices.
By Huayang Li, Tianyu Zhao, Deng Cai, Richard Sproat
arXiv:2509. 10534v3 Announce Type: replace-cross Abstract: The attention mechanism in a Transformer architecture matches key to query based on both content -- the what -- and position in a sequence -- the where.
By Anand Gopalakrishnan, Robert Csord\'as, J\"urgen Schmidhuber, Michael C. Mozer
The paper introduces PAYN, a training‑free token compression strategy for Multimodal Large Language Model (MLLM) based Referring Expression Segmentation (RES). By preserving original position embeddings and local spatial structures, PAYN retains tokens that are evenly distributed across neighboring regions, thereby maintaining spatial relational consistency. Experiments on multiple RES benchmarks show that PAYN outperforms existing token compression methods, confirming that position information alone is sufficient for effective compression in this task.
By Yuhan Liu, Yixiong Zou, Yuhua Li, Ruixuan Li
arXiv:2605. 25475v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference.
By Xintong Yang, Hao Gu, Binxing Xu, Lujun Li, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Yike Guo, Sirui Han
arXiv:2606. 24975v1 Announce Type: new Abstract: PaTH Attention showed that replacing RoPE's position-indexed rotations with accumulated data-dependent Householder reflections yields strong length extrapolation, though performance degrades at extreme context lengths.
By Mahesh Godavarti
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
By Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan
arXiv:2604. 00004v2 Announce Type: replace-cross Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT).
By Ning Yang, Hengyu Zhong, Wentao Wang, Baoliang Tian, Haijun Zhang, Jun Wang