arXiv:2607. 02509v1 Announce Type: new Abstract: Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications.
By Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He
CoEM introduces a Commit-on-Evidence Memory system that learns when to compress source evidence into compact memory facts while preserving potentially useful excerpts verbatim in a pending set. The system uses a learned policy to decide whether to promote, retain, or discard each pending excerpt as new context arrives, and a frozen verifier ensures only supported facts are committed. Reinforcement learning trains this policy with step-level evidence rewards and final answer rewards, leading to consistent improvements in long-context reasoning, achieving 10.4–11.4 F1 points over the strongest baseline on 6,400-document inputs.
By Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li, Can Rong, Heye Huang
arXiv:2609.07093v2 Announce Type: replace
Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...
By Yifan Wang, Xinkui Lin, Yongxiu Xu, Shen Gao, Ruochen Yang, Kun Huang, Yubin Wang, Jie Wu, Wei Liu, Jian Luan, Hongbo Xu, Shuo Shang
The paper introduces Highlight-Then-Summarize (H2S), a two-step approach that first highlights question-relevant evidence in long documents and then condenses it into a compact, question-conditioned summary before generating an answer. The authors built the H2S-Dataset with 6,647 examples spanning 11 benchmark families, and developed H2S-RL to reward evidence selection and summary construction. Evaluated on the H2S-Bench suite, the H2S-14B model outperforms larger open-source models, achieving the highest overall score and maintaining strong performance even with a reduced output budget.
By Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc), Gaofeng Lu (Baidu Inc), Zhenyu Jiao (Baidu Inc), Xing Yuan (Baidu Inc), Dai Dai (Baidu Inc), Tong Mo (Peking University), Long Zeng (Tsinghua University)
arXiv:2607. 19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier.
By Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang
arXiv:2606. 01223v1 Announce Type: cross Abstract: Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cues into high-level interpretations.
By Jingjie Lin, Bingbing Wang, Zihan Wang, Zhengda Jin, Weiming Qiao, Jing Li, Ruifeng Xu