arXiv:2606. 30573v1 Announce Type: new Abstract: We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks.
By Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He
arXiv:2606.21804v2 Announce Type: replace-cross
Abstract: Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding age...
By Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan, He He, Valerie Chen
Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses.
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation.
The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.
By Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan
arXiv:2608. 11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance.
By Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang