arXiv:2608.21833v1 Announce Type: new
Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especial...
By Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng
A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.
By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park
arXiv:2609.21293v1 Announce Type: new
Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessari...
By Xiuhui Zhang, Yi Chen, Shusheng Xu, Fan Li, Huan Wang, Tongkai Yang, Binhang Yuan
arXiv:2609.22000v1 Announce Type: new
Abstract: Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Re...
By Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou
arXiv:2610.08621v1 Announce Type: new
Abstract: Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experi...
By Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu
SWE-Game is a benchmark comprising 247 tasks based on 41 Godot games across 13 gameplay categories, testing coding agents on tasks such as brief-to-game, design-document implementation, skeleton completion, fault repair, and Godot-to-Unity porting. Evaluation uses engine-state checks, replay of certified reference inputs, and agent-authored demonstrations to judge mechanic correctness, playability, and post-repair behavior, supplemented by vision‑language rubrics for presentation. Across six models, Opus5 leads but overall scores stay below 60/100, highlighting common issues like omitted requirements and gameplay logic errors.
By Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang
arXiv:2605. 25160v2 Announce Type: replace Abstract: GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments.
By Guohong Liu, Jialei Ye, Pengzhi Gao, Wei Liu, Jian Luan, Yunxin Liu, Yuanchun Li
The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.
By Joey Xiao, Haonan Huang
arXiv:2609.09059v1 Announce Type: new
Abstract: While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existi...
By Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie
arXiv:2609.21562v1 Announce Type: cross
Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...
By Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu
arXiv:2606. 29957v1 Announce Type: cross Abstract: Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code.
By Yifan Wu, Zhuokai Zhao, Songlin Li, Ho Hin Lee, Jiacheng Zhu, Shirley Wu, Tianhe Yu, Serena Li, Lizhu Zhang, Xiangjun Fan, Shengzhi Li
GameGo is a framework that converts short game ideas into detailed Product Requirements Documents using industry practices, enabling coding agents to generate complete games from sparse user queries. It employs dynamic compression to keep essential gameplay constraints while allowing design flexibility. The authors built GameGoData with over 55,000 development trajectories and GameGoBench with 124 game queries, training GameGoCoder to outperform baselines and match leading models on gamedev benchmarks.
By Haoyue Yang, Jingyao Li, Zhengfan Wu, Jing Liu, Xuanle Zhao, Kang Liu