arXiv:2608.21833v1 Announce Type: new
Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especial...
By Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng
The paper introduces GUI agents for continual game generation, presenting PlaytestArena—a benchmark of 200 browser-based game-generation tasks with rubrics for in‑play behavior—and Play2Code, an iterative framework where a game agent and a rubric‑blind GUI playtester refine games through shared memory. Play2Code achieves a 66.8% rubric pass rate, surpassing baseline methods by 37.1 and 14.6 points, and shows consistent score improvement across refinement rounds. The study demonstrates that GUI playtesting provides actionable, traceable feedback that can guide interactive code generation.
By Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo
SWE-Game is a benchmark comprising 247 tasks based on 41 Godot games across 13 gameplay categories, testing coding agents on tasks such as brief-to-game, design-document implementation, skeleton completion, fault repair, and Godot-to-Unity porting. Evaluation uses engine-state checks, replay of certified reference inputs, and agent-authored demonstrations to judge mechanic correctness, playability, and post-repair behavior, supplemented by vision‑language rubrics for presentation. Across six models, Opus5 leads but overall scores stay below 60/100, highlighting common issues like omitted requirements and gameplay logic errors.
By Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang
A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.
By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park
GameGo is a framework that converts short game ideas into detailed Product Requirements Documents using industry practices, enabling coding agents to generate complete games from sparse user queries. It employs dynamic compression to keep essential gameplay constraints while allowing design flexibility. The authors built GameGoData with over 55,000 development trajectories and GameGoBench with 124 game queries, training GameGoCoder to outperform baselines and match leading models on gamedev benchmarks.
By Haoyue Yang, Jingyao Li, Zhengfan Wu, Jing Liu, Xuanle Zhao, Kang Liu
arXiv:2609.39045v1 Announce Type: new
Abstract: Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a pla...
By Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou, Siqi Liu, Aayush Salvi, Yiheng Lin, Ce Zhang, Xiaohan Lan, Jiahui Zhu, Yujie Zhong, Qi She, Biwei Huang
arXiv:2609.22308v1 Announce Type: new
Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing researc...
By Boyu Qiao, Zixin Tang, Xiaoshuai Hao, Wenbo Li
The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.
By Maxim Gumin, Hsueh-Ti Derek Liu, Victor Zordan, Daniel Ritchie
arXiv:2609.25652v1 Announce Type: new
Abstract: Recent game world models support realistic visual simulation and interactive gameplay based on player inputs. However, they typically learn environment...
By Zijun Lin, Zhiyang Deng, Yuzhe Wu, Bihan Wen, Yeying Jin
arXiv:2609.21293v1 Announce Type: new
Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessari...
By Xiuhui Zhang, Yi Chen, Shusheng Xu, Fan Li, Huan Wang, Tongkai Yang, Binhang Yuan
arXiv:2602. 11103v2 Announce Type: replace Abstract: Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind.
By Wayne Chi, Yixiong Fang, Arnav Yayavaram, Siddharth Yayavaram, Seth Karten, Qiuhong Anna Wei, Runkun Chen, Alexander Wang, Valerie Chen, Ameet Talwalkar, Chris Donahue
arXiv:2609.21562v1 Announce Type: cross
Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...
By Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu