The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv:2510. 07884v2 Announce Type: replace-cross Abstract: Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling.
By Houcheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang, Chen Gao, Xiang Wang, Xiangnan He, Yang Deng
arXiv:2606. 07604v1 Announce Type: cross Abstract: Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs).
By Harry Jake Cunningham, Nicola Muca Cirone
arXiv:2601. 04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias.
By Maryam Rahimi, Mahdi Nouri, Yadollah Yaghoobzadeh
arXiv:2609.37879v1 Announce Type: cross
Abstract: How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention....
By Timur Mudarisov, Mikhail Burtsev, Radu State
The paper investigates how hybrid attention mechanisms—combining full softmax attention with recurrent alternatives—affect multilingual language models, especially for long sequences and poorly tokenized languages. Interpretability analysis reveals that cross‑lingual representations form patterns linked to the ordering of recurrent and full‑attention layers, with a notable spike in alignment after the first full‑attention layer. Distillation experiments show that alternative layer orderings consistently outperform the standard arrangement, achieving up to 2.5× faster learning, suggesting that starting with a full‑attention layer may benefit multilingual models.