arXiv AI

Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

The paper investigates cost-inefficient behaviors in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent on SWE-bench Verified. It identifies three main inefficiencies—subsumed retrieval, similar script generation, and test re-execution—that affect 79–98% of tasks and contribute up to 22.75% of costs. The study evaluates mitigation strategies, finding that developer-designed skills reduce costs by up to 41.73%, outperforming other approaches.

arXiv AI
6d ago

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.

By Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Sep 7

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

The paper examines how AI coding agents are evolving beyond simple autocomplete to perform complex tasks such as repository inspection, multi-file editing, tool execution, test writing, pull request creation, and long-duration work with minimal supervision. It highlights that while these agents boost coding activity, significant bottlenecks remain in review, integration, testing, security, deployment, and production operations, and that the economics of software development are shifting toward variable token, tool, sandbox, CI, and rework costs. The authors synthesize recent research and industry data to propose four engineering concepts—Agentic SDLC Throughput Paradox, Production-Qualified Change, Verification Tax, and an Agentic SDLC Control Plane—to guide the allocation of autonomy within cost, reliability, and human-attention constraints, ultimately reframing the research focus to production-qualified value per dollar, reviewer-hour, and operational risk.

By Happy Bhati
arXiv AI
Sep 21

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

arXiv:2609.22068v1 Announce Type: new Abstract: Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich sourc...

By Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo