LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
Read the original on arXiv Computation and Language →The paper introduces LongHarness Bench, a new benchmark designed to evaluate both the effectiveness and efficiency of language model harnesses for long-context reasoning. It features tasks that require diverse retrieval strategies—such as lexical search and semantic matching—and strategic, adaptive reasoning over global and local context, with only a small subset of the context being useful at each step. Evaluations across multiple state‑of‑the‑art models and harnesses show that even strong combinations achieve only 68% macro‑average accuracy, and that the same model can vary markedly in efficiency depending on the harness used.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.