arXiv AI By Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo

GUI Agents for Continual Game Generation

Read the original on arXiv AI →

The paper introduces GUI agents for continual game generation, presenting PlaytestArena—a benchmark of 200 browser-based game-generation tasks with rubrics for in‑play behavior—and Play2Code, an iterative framework where a game agent and a rubric‑blind GUI playtester refine games through shared memory. Play2Code achieves a 66.8% rubric pass rate, surpassing baseline methods by 37.1 and 14.6 points, and shows consistent score improvement across refinement rounds. The study demonstrates that GUI playtesting provides actionable, traceable feedback that can guide interactive code generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Oct 1

A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.

By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park
arXiv Computation and Language
Sep 21

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

arXiv:2609.22000v1 Announce Type: new Abstract: Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Re...

By Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Keliang Li, Ning Li, Wanli Li, Dayiheng Liu, Dunjie Lu, Changwei Luo, Que Shen, Zheyuan Wang, Zijian Wang, Jie Wu, Gao Wu, Zhihui Xie, Rui Xie, Haiyang Xu, An Yang, Jiakang Yuan, Yanming Zhang, Jiajun Zhang, Xi Zhang, Zhenru Zhang, Zhuo Zhen, Mingkang Zhu, Bowen Zhou
arXiv AI
1d ago

SWE-Game: Can Coding Agents Build the Games We Want?

SWE-Game is a benchmark comprising 247 tasks based on 41 Godot games across 13 gameplay categories, testing coding agents on tasks such as brief-to-game, design-document implementation, skeleton completion, fault repair, and Godot-to-Unity porting. Evaluation uses engine-state checks, replay of certified reference inputs, and agent-authored demonstrations to judge mechanic correctness, playability, and post-repair behavior, supplemented by vision‑language rubrics for presentation. Across six models, Opus5 leads but overall scores stay below 60/100, highlighting common issues like omitted requirements and gameplay logic errors.

By Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang