arXiv:2607. 08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them.
By Siddhartha Jain, Ameya Velingker
The paper introduces LongHarness Bench, a new benchmark designed to evaluate both the effectiveness and efficiency of language model harnesses for long-context reasoning. It features tasks that require diverse retrieval strategies—such as lexical search and semantic matching—and strategic, adaptive reasoning over global and local context, with only a small subset of the context being useful at each step. Evaluations across multiple state‑of‑the‑art models and harnesses show that even strong combinations achieve only 68% macro‑average accuracy, and that the same model can vary markedly in efficiency depending on the harness used.
By Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye
arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.
By Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
arXiv:2505. 19293v2 Announce Type: replace-cross Abstract: Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhausting tasks -- e.
By Wang Yang, Hongye Jin, Shaochen Zhong, Song Jiang, Qifan Wang, Vipin Chaudhary, Xiaotian Han
arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.
By Purbesh Mitra, Sennur Ulukus
Randomized YaRN is a training method that enhances length generalization for large language models by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum. During training on short-context data, tokens receive YaRN positional encodings sampled from a larger position range, exposing the model to out-of-distribution positional representations. Evaluated on BABILong, Multi-Round Coreference Resolution, and LongBench v2, Randomized YaRN consistently improves reasoning performance on context lengths from 16K to 128K, outperforming standard fine‑tuning especially at far out‑of‑distribution lengths.
By Manas Mehta, Fangcong Yin, Greg Durrett
The paper introduces Highlight-Then-Summarize (H2S), a two-step approach that first highlights question-relevant evidence in long documents and then condenses it into a compact, question-conditioned summary before generating an answer. The authors built the H2S-Dataset with 6,647 examples spanning 11 benchmark families, and developed H2S-RL to reward evidence selection and summary construction. Evaluated on the H2S-Bench suite, the H2S-14B model outperforms larger open-source models, achieving the highest overall score and maintaining strong performance even with a reduced output budget.
By Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc), Gaofeng Lu (Baidu Inc), Zhenyu Jiao (Baidu Inc), Xing Yuan (Baidu Inc), Dai Dai (Baidu Inc), Tong Mo (Peking University), Long Zeng (Tsinghua University)
arXiv:2608. 03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length.
By Dawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou
arXiv:2603. 02112v2 Announce Type: replace Abstract: Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning.
By Chenxiao Yang, Nathan Srebro, Zhiyuan Li
arXiv:2505. 17315v2 Announce Type: replace Abstract: Recent language models exhibit strong reasoning capabilities, yet the influence of long-context capacity on reasoning remains underexplored.
By Wang Yang, Zirui Liu, Hongye Jin, Qingyu Yin, Vipin Chaudhary, Xiaotian Han
arXiv:2505.23126v5 Announce Type: replace
Abstract: Although many benchmarks evaluate the reasoning abilities of Large Language Models (LLMs) within domains such as mathematics, coding, or data wrang...
By Atharva Naik, Prakam, Yash Mathur, Darsh Agrawal, Manav Kapadnis, Yuwei An, Clayton Marr, Carolyn Rose, David Mortensen
arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.
By Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong