arXiv:2609.22101v1 Announce Type: cross
Abstract: Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confu...
By Meysam Ghaffari, Nina Fatehi, Bhaskar Sen, Nasim Sabetpour, Carlos Morato
arXiv:2607. 19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier.
By Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang
The paper investigates how large language models acquire knowledge during pre‑training, proposing that auxiliary views—reformulations of knowledge—are causally beneficial. Experiments show that repetition is essential, paraphrasing helps only at smaller batch sizes, and reallocating tokens from repetition to auxiliary views improves learning even for factual recall. The study also finds that the benefit of auxiliary views does not depend on the teacher model’s strength, identifies specific knowledge types that aid learning, and explores mechanistic effects via layer‑wise biases and compression.
By Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen
arXiv:2607. 02509v1 Announce Type: new Abstract: Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications.
By Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He
arXiv:2510. 01163v2 Announce Type: replace Abstract: The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples.
By Wa\"iss Azizian, Ali Hasan
arXiv:2606. 09525v1 Announce Type: cross Abstract: During instruction fine-tuning (IFT), large language models (LLMs) learn to follow instructions by using the provided context to answer a query.
By Nadya Yuki Wangsajaya, Haeun Yu, Isabelle Augenstein
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization.
arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.
By Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
MoSAR introduces a mixture of semantic attention regimes that learns an adaptive, distance‑dependent attention geometry from data, rather than predefining sparse or local patterns. The model uses input‑conditioned routers to select short, medium, or global regimes, creating a continuous attention field that can be discretized for efficient inference. Experiments show that MoSAR achieves lower‑reach attention without sacrificing language‑modeling quality, improving perplexity over dense RoPE and outperforming baselines like ALiBi, while remaining stable under top‑1 discretization.
By Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
arXiv:2505. 17315v2 Announce Type: replace Abstract: Recent language models exhibit strong reasoning capabilities, yet the influence of long-context capacity on reasoning remains underexplored.
By Wang Yang, Zirui Liu, Hongye Jin, Qingyu Yin, Vipin Chaudhary, Xiaotian Han
The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv:2507. 05019v2 Announce Type: replace-cross Abstract: In-context learning enables transformer models to generalize to new tasks based solely on input prompts, without any need for weight updates.
By Lorenzo Braccaioli, Anna Vettoruzzo, Prabhant Singh, Joaquin Vanschoren, Mohamed-Rafik Bouguelia, Nicola Conci