arXiv:2609.22101v1 Announce Type: cross
Abstract: Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confu...
By Meysam Ghaffari, Nina Fatehi, Bhaskar Sen, Nasim Sabetpour, Carlos Morato
arXiv:2607. 19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier.
By Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang
The paper investigates how large language models acquire knowledge during pre‑training, proposing that auxiliary views—reformulations of knowledge—are causally beneficial. Experiments show that repetition is essential, paraphrasing helps only at smaller batch sizes, and reallocating tokens from repetition to auxiliary views improves learning even for factual recall. The study also finds that the benefit of auxiliary views does not depend on the teacher model’s strength, identifies specific knowledge types that aid learning, and explores mechanistic effects via layer‑wise biases and compression.
By Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen
arXiv:2607. 02509v1 Announce Type: new Abstract: Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications.
By Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He
arXiv:2510. 01163v2 Announce Type: replace Abstract: The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to adapt to new tasks from only a handful of examples.
By Wa\"iss Azizian, Ali Hasan
arXiv:2606. 09525v1 Announce Type: cross Abstract: During instruction fine-tuning (IFT), large language models (LLMs) learn to follow instructions by using the provided context to answer a query.
By Nadya Yuki Wangsajaya, Haeun Yu, Isabelle Augenstein