arXiv AI By Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, Tieyong Zeng

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Read the original on arXiv AI →

The paper introduces a controlled evaluation to disentangle answer coverage, repeatable task advantages, and gains from pre‑execution selection in large‑language‑model (LLM) harnesses. On 386 MATH‑500 tasks, eight generated harnesses and a baseline with nine identical copies were compared over three executions each, revealing that identical programs provide a 2.16‑point repeat‑averaged oracle headroom while generated programs show more repeatable score patterns but mainly expose persistent weaknesses. The study concludes that coverage and repeatability alone cannot justify claims of useful specialization and proposes an evaluation standard for harness diversity that requires task advantages to persist across executions and improve on additional fixed‑program executions under matched inference budgets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati