Coaching Qwen3 Coder 30B to Think Like a CodeClash Arena Agent
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Large language model coding agents have recently become useful for software tasks, but weaker or open-weight agents still struggle to reliably interpret user intent and execute complex multi-step work...
The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.
arXiv:2606. 30573v1 Announce Type: new Abstract: We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks.
arXiv:2609.27717v1 Announce Type: new Abstract: Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather...
arXiv:2606. 10933v1 Announce Type: new Abstract: LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories.
arXiv:2609.22068v1 Announce Type: new Abstract: Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich sourc...