Prompt-Induced Waste in Coding Agents: Reasoning, Effort, Harness Design, and End-to-End Cost
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses.
The study investigates how individual components of a coding harness—planning, action space, and context management—affect autonomous coding agents’ performance. By fixing the execution loop and varying these components across 176 settings on SWE‑Bench Verified and Terminal‑Bench 2.1, the authors find that context management is most valuable when context windows are tight, staging rule‑based elision before LLM summarization yields the best efficiency, planning serves as an accuracy scaffold for weaker models and a cost saver for stronger ones, and predefined tools help models with limited bash skills while bash‑capable models benefit from a bash‑only interface. Trajectory‑level analysis shows that context management lengthens execution paths, planning alters where trajectories terminate, and the action space determines code granularity, offering a modular framework for future harness design.
arXiv:2603.19896v2 Announce Type: replace Abstract: Tool-using large language model (LLM) agents often face a fundamental tension between answer quality and execution cost. Fixed workflows are stable...
arXiv:2609.01437v1 Announce Type: cross Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly...
The paper introduces Growing Harness, a training method that transforms recurring control logic in large language model agents into reusable executable code, reducing reliance on the model for task-specific decisions. By using strategy-free scaffolds, failure-guided code repair, and success-first gating, the approach learns a shared harness that improves performance across multiple benchmarks and model sizes. Experiments on BrowseComp-Plus and WebArena-Verified show significant gains in success rates and substantial reductions in LLM calls and inference cost compared to traditional tool‑calling agents.
The paper investigates cost-inefficient behaviors in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent on SWE-bench Verified. It identifies three main inefficiencies—subsumed retrieval, similar script generation, and test re-execution—that affect 79–98% of tasks and contribute up to 22.75% of costs. The study evaluates mitigation strategies, finding that developer-designed skills reduce costs by up to 41.73%, outperforming other approaches.