Introducing HELMET: Holistically Evaluating Long-context Language Models
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2610.02404v1 Announce Type: new Abstract: We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simple...
The paper introduces LongHarness Bench, a new benchmark designed to evaluate both the effectiveness and efficiency of language model harnesses for long-context reasoning. It features tasks that require diverse retrieval strategies—such as lexical search and semantic matching—and strategic, adaptive reasoning over global and local context, with only a small subset of the context being useful at each step. Evaluations across multiple state‑of‑the‑art models and harnesses show that even strong combinations achieve only 68% macro‑average accuracy, and that the same model can vary markedly in efficiency depending on the harness used.
arXiv:2603. 20843v3 Announce Type: replace-cross Abstract: Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information structuring remains largely implicit in existing approaches.
arXiv:2609.22452v1 Announce Type: new Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness.
arXiv:2505. 19293v2 Announce Type: replace-cross Abstract: Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhausting tasks -- e.