arXiv AI

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.

Hugging Face Trending Papers
Jul 9

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness.

arXiv AI
6d ago

Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding

The paper introduces Highlight-Then-Summarize (H2S), a two-step approach that first highlights question-relevant evidence in long documents and then condenses it into a compact, question-conditioned summary before generating an answer. The authors built the H2S-Dataset with 6,647 examples spanning 11 benchmark families, and developed H2S-RL to reward evidence selection and summary construction. Evaluated on the H2S-Bench suite, the H2S-14B model outperforms larger open-source models, achieving the highest overall score and maintaining strong performance even with a reduced output budget.

By Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc), Gaofeng Lu (Baidu Inc), Zhenyu Jiao (Baidu Inc), Xing Yuan (Baidu Inc), Dai Dai (Baidu Inc), Tong Mo (Peking University), Long Zeng (Tsinghua University)
arXiv Computation and Language
Sep 11

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

The paper proposes a new method for training large language models to handle long-context reasoning by combining Group Relative Policy Optimization (GRPO) with on‑policy distillation (OPD). It introduces a synthetic multilingual dataset called LongBlocks that tests multi‑hop reasoning, contextual grounding, and long‑form generation. Experiments show that the combined approach outperforms either GRPO or OPD alone while maintaining short‑context performance.

By Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins
Hugging Face Trending Papers
Jul 2

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization.

arXiv AI
Aug 6

Chained Recursive Language Models for Multi-Iteration Reasoning

arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.

By Purbesh Mitra, Sennur Ulukus
arXiv AI
4d ago

CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory

CoEM introduces a Commit-on-Evidence Memory system that learns when to compress source evidence into compact memory facts while preserving potentially useful excerpts verbatim in a pending set. The system uses a learned policy to decide whether to promote, retain, or discard each pending excerpt as new context arrives, and a frozen verifier ensures only supported facts are committed. Reinforcement learning trains this policy with step-level evidence rewards and final answer rewards, leading to consistent improvements in long-context reasoning, achieving 10.4–11.4 F1 points over the strongest baseline on 6,400-document inputs.

By Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li, Can Rong, Heye Huang
arXiv Computation and Language
4d ago

Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs

The paper proposes Layer-Informed Fine-Tuning (LIFT), a method that identifies and updates only the most functionally critical layers of large language models (LLMs) using a bottleneck identification mechanism based on sensitivity analysis. By focusing on layers that handle conceptualization, reasoning, and textualization, LIFT aims to accelerate training and enhance performance on reasoning tasks. Experiments demonstrate that this selective fine-tuning approach both speeds up the training process and yields significant performance gains.

By Junning Shao, Siwei Wang, Zhixuan Fang