arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.
By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv:2607. 05297v1 Announce Type: new Abstract: Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability.
By Zefeng Wang, Minxi Yan, Jinhe Bi, Sikuan Yan, Volker Tresp, Yunpu Ma
arXiv:2607. 21596v2 Announce Type: replace Abstract: Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution.
By Zeyu Ren, Ling Yue, Ran Li, Yishu Wang, Shengxiang Xu, Hanmo Liu, Shaowu Pan, Shimin Di
arXiv:2608. 08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment.
By Tailin Zhou
arXiv:2603.01209v3 Announce Type: replace
Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python v...
By Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen
SkillEffect is a checked‑lowering runtime that ensures agent tool calls stay within memory limits by verifying each proposed program against an immutable input before execution. It uses audited relation plugins to provide source recognition, bounded intermediate representation construction, and postconditions, while a shared runtime handles selection, bounded VM execution, and atomic capacity leasing. Experiments across six operator families show that bounded access significantly reduces peak memory usage and improves completion rates under fixed memory caps.
By Yinuo Wang, Yiyu Shi
arXiv:2608. 02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups.
By Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies)
The paper introduces VERSE, a Verified Self‑Evolving optimizer that enhances LLM agent harnesses by allowing the optimizer to test edits, replay failures, and perturb steps while tracking fixes and regressions. VERSE builds its own tools for failure analysis, verification, training audits, and workflow control, and uses this feedback to revise the harness’s prompts, skills, tools, hooks, and notes without changing model weights. In experiments across five executors and multiple languages, VERSE improves all evaluated harness optimizers, achieving higher accuracy on held‑out and out‑of‑distribution tasks compared to the strongest baselines.
By Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.
By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
arXiv:2607. 20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally.
By Zibin Lin, Shengli Zhang, Taotao Wang, Yihan Xia, Deen Ma, Guofu Liao
The paper introduces InFlowOp, a label‑free optimization framework that assigns costs to each decision in a multi‑agent workflow, balancing agent competence against execution time. It determines task granularity and agent assignment before execution and corrects faults during execution using the same cost metric. The authors also present Braid, a benchmark for multi‑agent coordination, and show that InFlowOp outperforms single‑agent baselines by up to 11.97% across various domains.
By Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng, Qingyun Wang, Haifeng Chen
Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing too...