arXiv:2608.01347v4 Announce Type: replace
Abstract: Coding-agent efficiency cannot be characterized by token count or model price alone. We study how end-to-end cost and task success depend jointly o...
By Sarel Weinberger, Amir Hozez
The paper introduces agentic meta‑reasoning, a structured inference‑time framework that explicitly manages control decisions—such as selecting partial work, restarting, or stopping—during long‑horizon agentic tasks. By delegating task execution to workers and consolidating decisions through a lightweight controller that references persistent memory, the method reduces the need to replay full histories. Experiments on ProgramBench and other benchmarks show that meta‑reasoning improves performance over direct control baselines, especially as computation budgets increase, and reveals greater reuse of earlier work and higher solution coverage.
By Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal
arXiv:2608. 07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain.
By Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani
arXiv:2607. 02436v1 Announce Type: cross Abstract: Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the assumption that more capability yields better software.
By Achint Mehta
The paper investigates how altering the harness—specifically the way a coding agent manages context and tool outputs—affects performance when the underlying model and task remain unchanged. Two harness configurations were compared on three coding benchmarks: a control that preserves the full conversation in order, and a treatment that mechanically shortens older tool results to keep the context tight. Across all benchmarks, the treatment increased the mean per‑task fail‑to‑pass fraction and, in some cases, the number of complete solutions, demonstrating that the harness itself can significantly influence a frozen model’s effectiveness.
By Sydney Lewis
The paper proposes a unified framework for test‑time reasoning methods, framing them as recursion operators—GROW, PRUNE, and BRANCH—applied to an agent’s reasoning trace. Experiments across five benchmarks and three frontier models show that BRANCH, which samples and selects among multiple reasoning paths, consistently outperforms the other operators and a single‑pass chain‑of‑thought baseline, improving accuracy by an average of 5.98 percentage points. The study also highlights the importance of paired evaluation and careful handling of scoring‑pipeline failures, as these factors can significantly alter comparative outcomes.
By Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie
arXiv:2606. 30573v1 Announce Type: new Abstract: We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks.
By Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He
The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.
By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.
By Md Jueal Mia, M. Hadi Amini
arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.
By Peijia Qin, Qi Cao, Pengtao Xie
arXiv:2607. 13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires.
By Junjie Yin, Xinyu Feng
arXiv:2609.16302v1 Announce Type: cross
Abstract: When a coding agent returns to existing software, it inherits evidence from earlier engineering work: tests, type checks, proofs, static analyses, an...
By Anjan Goswami