arXiv AI

Scaffold Effects on GAIA: A Controlled Comparison

arXiv:2606. 08529v1 Announce Type: new Abstract: Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well characterized under controlled conditions.

arXiv AI
Sep 3

Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization

The paper introduces Belief-Calibrated Optimization (BCO), a method that records and updates a persistent in‑context document representing an agent’s belief about how the environment responds to edits. By continually revising this world model as new candidates are evaluated, BCO improves the performance of frozen LLM agents across five benchmarks, outperforming a control lacking the world model. An offline ablation shows that the document’s content provides reusable, accurate predictions of environmental responses, beyond mere form.

By Yuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia
arXiv AI
Sep 25

SheetMind: Actions Set Accuracy, Agents Set the Failure Mode

SheetMind is a Manager‑Action‑Reflection framework that evaluates how much spreadsheet agent performance derives from the agents themselves versus the shared action interface. In a controlled study on all 221 tasks of the SheetCopilot Benchmark, replacing the high‑level action API with primitive cell operations drops accuracy by 47.1 points, while adding a Reflection Agent improves performance by 4.5 points and a Manager by 1.4 points. The framework also shows that decomposition changes failure modes, reducing silent wrong outputs from 33% to 25%, and that GPT‑5 and GPT‑5‑mini achieve similar performance, whereas GPT‑3.5 underperforms significantly.

By Lyuhao Chen, Xi Cheng, Yanming Kang, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Yulang Fei, Brian Zhu, Daniel Jin, Binze Cai, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao
arXiv AI
Aug 11

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

arXiv:2608. 08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tools.

By Marc Alier Forment, Mar\'ia Jos\'e Casa\~n Guerrero, Francisco Jos\'e Garc\'ia-Pe\~nalvo, Juanan Pereira
arXiv AI
Aug 28

Same Model, Different Harness: Different Coding-Agent Results

The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.

By Sydney Lewis