CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.
arXiv:2609.21562v1 Announce Type: cross Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...
arXiv:2609.21293v1 Announce Type: new Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessari...
BuildBench introduces a realistic benchmark for evaluating large language model agents on the task of compiling open‑source software (OSS). It includes diverse OSS projects that lack clear build instructions, have undocumented dependencies, and may require source patching or script modification. The authors also present OSS‑BUILD‑AGENT, a baseline LLM‑based agent that retrieves build instructions effectively and achieves state‑of‑the‑art performance on the benchmark.
arXiv:2608.21833v1 Announce Type: new Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especial...
arXiv:2606. 07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.