Hugging Face Trending Papers

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Read the original on Hugging Face Trending Papers →

Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
4d ago

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

The paper introduces LongHarness Bench, a new benchmark designed to evaluate both the effectiveness and efficiency of language model harnesses for long-context reasoning. It features tasks that require diverse retrieval strategies—such as lexical search and semantic matching—and strategic, adaptive reasoning over global and local context, with only a small subset of the context being useful at each step. Evaluations across multiple state‑of‑the‑art models and harnesses show that even strong combinations achieve only 68% macro‑average accuracy, and that the same model can vary markedly in efficiency depending on the harness used.

By Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye
arXiv AI
Jul 8

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.

By Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
arXiv AI
Aug 6

Chained Recursive Language Models for Multi-Iteration Reasoning

arXiv:2608. 05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer.

By Purbesh Mitra, Sennur Ulukus
arXiv Computation and Language
Sep 1

Randomized YaRN Improves Length Generalization for Long-Context Reasoning

Randomized YaRN is a training method that enhances length generalization for large language models by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum. During training on short-context data, tokens receive YaRN positional encodings sampled from a larger position range, exposing the model to out-of-distribution positional representations. Evaluated on BABILong, Multi-Round Coreference Resolution, and LongBench v2, Randomized YaRN consistently improves reasoning performance on context lengths from 16K to 128K, outperforming standard fine‑tuning especially at far out‑of‑distribution lengths.

By Manas Mehta, Fangcong Yin, Greg Durrett