arXiv AI By Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu, Lin Tan

Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

Read the original on arXiv AI →

The paper investigates cost-inefficient behaviors in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent on SWE-bench Verified. It identifies three main inefficiencies—subsumed retrieval, similar script generation, and test re-execution—that affect 79–98% of tasks and contribute up to 22.75% of costs. The study evaluates mitigation strategies, finding that developer-designed skills reduce costs by up to 41.73%, outperforming other approaches.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.

By Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan