arXiv AI
3d ago

GUI Agents for Continual Game Generation

The paper introduces GUI agents for continual game generation, presenting PlaytestArena—a benchmark of 200 browser-based game-generation tasks with rubrics for in‑play behavior—and Play2Code, an iterative framework where a game agent and a rubric‑blind GUI playtester refine games through shared memory. Play2Code achieves a 66.8% rubric pass rate, surpassing baseline methods by 37.1 and 14.6 points, and shows consistent score improvement across refinement rounds. The study demonstrates that GUI playtesting provides actionable, traceable feedback that can guide interactive code generation.

By Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo
arXiv AI
1d ago

SWE-Game: Can Coding Agents Build the Games We Want?

SWE-Game is a benchmark comprising 247 tasks based on 41 Godot games across 13 gameplay categories, testing coding agents on tasks such as brief-to-game, design-document implementation, skeleton completion, fault repair, and Godot-to-Unity porting. Evaluation uses engine-state checks, replay of certified reference inputs, and agent-authored demonstrations to judge mechanic correctness, playability, and post-repair behavior, supplemented by vision‑language rubrics for presentation. Across six models, Opus5 leads but overall scores stay below 60/100, highlighting common issues like omitted requirements and gameplay logic errors.

By Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang
arXiv AI
Oct 1

A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.

By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park
arXiv AI
1d ago

GAMEGO: Training Game-Dev Agents with Synthetic Trajectories Anchored in Real-World Assets

GameGo is a framework that converts short game ideas into detailed Product Requirements Documents using industry practices, enabling coding agents to generate complete games from sparse user queries. It employs dynamic compression to keep essential gameplay constraints while allowing design flexibility. The authors built GameGoData with over 55,000 development trajectories and GameGoBench with 124 game queries, training GameGoCoder to outperform baselines and match leading models on gamedev benchmarks.

By Haoyue Yang, Jingyao Li, Zhengfan Wu, Jing Liu, Xuanle Zhao, Kang Liu