arXiv Machine Learning

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

arXiv:2606. 20512v1 Announce Type: cross Abstract: LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong fixes) that does not exist in the code itself.

arXiv AI
Aug 28

Same Model, Different Harness: Different Coding-Agent Results

The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.

By Sydney Lewis
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv AI
3d ago

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

RankEvolve is an auto‑research framework that evolves generative ranking models by orchestrating multiple large‑language‑model coding agents through an Executable Operating Protocol (EOP). The system compiles a state machine that enforces phases, gates, branches, and loops, while a meta‑meta‑harness lets agents review and repair each other’s code. In budget‑matched experiments, heterogeneous composition of agents raised execution accuracy from 45.8 % to 62.5 % and reduced silent critical‑defect rates, achieving notable gains on the HSTU recommender and other benchmarks.

By Zheng Chen, Linfeng Liu, Hong Li, Hong Yan
arXiv Machine Learning
Sep 10

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

RubricRefine is a training‑free pre‑execution refinement method that generates task‑specific rubrics from tool documentation, scores candidate code against explicit contract checks, and iteratively repairs failures before execution. It achieves an average score of 0.86 across seven models on M3ToolEval without any execution attempts, outperforming prior inference‑time baselines while incurring lower latency. The approach shows consistent performance on single‑step API‑Bank tasks and maintains an advantage in multi‑turn settings on AppWorld, with its effectiveness tied to the quality of the supplied documentation.

By Will LeVine, Brendan Evers, Sam Saltwick, Abhay Venkatesh