arXiv AI

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

arXiv:2604. 00392v2 Announce Type: replace-cross Abstract: Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, and depend on.

arXiv AI
Aug 7

Recursive Synthesis for Long-Horizon Terminal Tasks

arXiv:2608. 05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent.

By Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
arXiv AI
Sep 2

REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.

By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
arXiv AI
2d ago

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Mingbird is a local‑first agent harness designed for small open‑weight language models (2–9 B) that run on ordinary laptops. It introduces ten mechanisms—such as a byte‑level net‑zero prefill budget, a finish gate that re‑reads the task before accepting completion, and signature‑level loop detection—to address common failure modes that arise from the harness rather than the model itself. In controlled experiments on the LRAB benchmark and the $ au^2$‑bench, Mingbird achieves higher overall scores (0.886 and 0.856 respectively) compared to other harnesses, and its ablation studies show that each mechanism contributes measurable performance gains.

By Hao Wang, Ting Huang
arXiv AI
3d ago

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

RankEvolve is an auto‑research framework that evolves generative ranking models by orchestrating multiple large‑language‑model coding agents through an Executable Operating Protocol (EOP). The system compiles a state machine that enforces phases, gates, branches, and loops, while a meta‑meta‑harness lets agents review and repair each other’s code. In budget‑matched experiments, heterogeneous composition of agents raised execution accuracy from 45.8 % to 62.5 % and reduced silent critical‑defect rates, achieving notable gains on the HSTU recommender and other benchmarks.

By Zheng Chen, Linfeng Liu, Hong Li, Hong Yan
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv AI
Jun 24

LemonHarness Technical Report

arXiv:2606. 24311v1 Announce Type: new Abstract: As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration.

By Kailong Ren, Fubo Sun, Jiachen Liu, Liu Yang, Zimo Yin, Jiaying Li, Congli Yin, Ming He, Yu Huo, Jiawei Liu, Zeping Chen, Yubin Huangfu, Ronghua Li, Yixuan Wu, Xing Su, Yanzhi Xu, Likang Wu, Hongke Zhao, Lei Zhang, Xiaohui Geng, Jianping Fan