arXiv AI By Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Read the original on arXiv AI →

The paper investigates the cost–quality trade-offs involved when coding agents switch between low‑cost, low‑capability (LC) and high‑cost, high‑capability (HC) language models during long‑running tasks. By experimenting with different handoff directions, timings, and interfaces—full‑trajectory transfer, compaction, and trajectory removal—the authors find that full‑trajectory escalation recovers less than half of the LC‑to‑HC quality gap while adding significant cost, a penalty they call the handoff tax. Conversely, downshifting from HC to LC yields a more favorable cost‑quality balance, and the optimal interface depends on the direction of the handoff. whyItMatters":"The study quantifies how model handoffs impact both performance and expense, offering guidance for designing more efficient coding agents that balance cost and quality."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 11

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

arXiv:2606. 12344v1 Announce Type: new Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.

By Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang