arXiv AI

Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks

arXiv:2608. 03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant information is preserved.

arXiv AI
Sep 1

MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects

MUDDLE is a benchmark designed to disentangle the effects of document length and topical distractors on document question‑answering systems. It contains 270 human‑annotated questions, each tested in five conditions: the source alone, the source with two or four hard negatives (topically similar), and the source with two or four random distractors matched in length and provenance. Experiments with GPT‑5‑mini show that hard negatives reduce accuracy more than length‑matched random distractors, indicating that topical similarity is a more significant source of error than length alone.

By Jason Luo, Saibilila Abudukelimu, Judy Song, Andrew Feng, Shivank Garg, Vasu Sharma, Kevin Zhu
arXiv Computation and Language
Sep 24

MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors

The paper introduces MWE‑ECL, a bilingual diagnostic framework that tests whether distant discourse anchors can override local lexical priors in multi‑word expression interpretation. It evaluates models on a 0‑128K context grid, finding that while retrieval of anchors is near perfect, the ability to change locally preferred readings varies, especially when the model’s default conflicts with the anchor. The study shows that explicit recoverability does not always translate into behavioral influence, with gaps differing across models and languages.

By Wei He, Aline Villavicencio, Rodrigo Wilkens, Zhenyun Deng
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos