More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
Read the original on arXiv AI →The paper introduces a controlled evaluation to disentangle answer coverage, repeatable task advantages, and gains from pre‑execution selection in large‑language‑model (LLM) harnesses. On 386 MATH‑500 tasks, eight generated harnesses and a baseline with nine identical copies were compared over three executions each, revealing that identical programs provide a 2.16‑point repeat‑averaged oracle headroom while generated programs show more repeatable score patterns but mainly expose persistent weaknesses. The study concludes that coverage and repeatability alone cannot justify claims of useful specialization and proposes an evaluation standard for harness diversity that requires task advantages to persist across executions and improve on additional fixed‑program executions under matched inference budgets.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.