Turbo Harness: Instance-Adaptive Harness Optimization
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
arXiv:2609.14857v1 Announce Type: new Abstract: Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve executio...
The paper introduces Harness Primitives—reusable agent harness mechanisms mined from failed task trajectories—and a framework called STITCH that selects and composes these primitives into task‑specific harnesses at test time. This approach avoids generating or debugging harness code for each task, achieving up to 12‑point gains in task success over fixed harness baselines and outperforming human‑designed harnesses like Codex CLI. STITCH also demonstrates minimal test‑time overhead (2.7%) and scales efficiently with the size of the primitive library.
arXiv:2606. 01770v1 Announce Type: cross Abstract: Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infrastructure from execution feedback, but they are typically evaluated on fixed offline benchmarks.
arXiv:2607. 12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents.
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.