arXiv:2608. 12218v1 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories.
By Arda Uzunoglu, Benjamin van Durme, Daniel Khashabi
The paper proposes a content‑based addressing scheme for long‑context models that replaces the growing token counter in Rotary Position Embedding (RoPE) with unit‑level addresses derived from the content of each unit. By dividing the token stream into units, the method preserves local RoPE behavior while allowing new units to be addressed via learned content maps, avoiding positional mismatches when extending context length. Experiments on character‑level Tiny Shakespeare show that a model trained on 256‑character contexts achieves lower perplexity at 4096 characters using this scheme, and a second diagnostic demonstrates retrieval of multiple serialized facts.
By Mahesh Godavarti
arXiv:2605. 28854v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL).
By Hua-Dong Xiong, Li Ji-An, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weaknes...
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.
By Suvadeep Hajra
arXiv:2609.39929v1 Announce Type: cross
Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and disting...
By Yuyang Wu, Yufeng Du, Hao Peng
arXiv:2510. 01163v2 Announce Type: replace Abstract: The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples.
By Wa\"iss Azizian, Ali Hasan
The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
By Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
The paper investigates distance generalization in transformer models, focusing on how well they can handle changes in inter-token distances between training and inference while keeping context length fixed. Using two synthetic delay-copy tasks that require copying tokens after finite delays, the authors evaluate the impact of positional encoding schemes (RoPE, ALiBi, and NoPE), the diversity of distances seen during training, and the conditions under which distance transfer learning is beneficial or detrimental. Their comprehensive study highlights the importance of understanding the underlying mechanisms that govern distance generalization in transformers.
By Daniel Henrik Nevermann, Claudius Gros
arXiv:2607. 19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.
By Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
By Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan
arXiv:2608.30315v1 Announce Type: new
Abstract: Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern lang...
By Junjie Yao, Liangkai Hang, Zhi-Qin John Xu