arXiv:2608.21833v1 Announce Type: new
Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especial...
By Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng
arXiv:2609.25001v1 Announce Type: new
Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
By Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
arXiv:2610.08621v1 Announce Type: new
Abstract: Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experi...
By Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu
arXiv:2609.09059v1 Announce Type: new
Abstract: While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existi...
By Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie
arXiv:2609.25652v1 Announce Type: new
Abstract: Recent game world models support realistic visual simulation and interactive gameplay based on player inputs. However, they typically learn environment...
By Zijun Lin, Zhiyang Deng, Yuzhe Wu, Bihan Wen, Yeying Jin
A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.
By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park
The paper introduces GUI agents for continual game generation, presenting PlaytestArena—a benchmark of 200 browser-based game-generation tasks with rubrics for in‑play behavior—and Play2Code, an iterative framework where a game agent and a rubric‑blind GUI playtester refine games through shared memory. Play2Code achieves a 66.8% rubric pass rate, surpassing baseline methods by 37.1 and 14.6 points, and shows consistent score improvement across refinement rounds. The study demonstrates that GUI playtesting provides actionable, traceable feedback that can guide interactive code generation.
By Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo
arXiv:2602. 11103v2 Announce Type: replace Abstract: Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind.
By Wayne Chi, Yixiong Fang, Arnav Yayavaram, Siddharth Yayavaram, Seth Karten, Qiuhong Anna Wei, Runkun Chen, Alexander Wang, Valerie Chen, Ameet Talwalkar, Chris Donahue
arXiv:2508.16821v2 Announce Type: replace
Abstract: We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinfo...
By Sam Earle, Graham Todd, Yuchen Li, Ahmed Khalifa, Muhammad Umair Nasir, Zehua Jiang, Andrzej Banburski-Fahey, Julian Togelius
arXiv:2608.24680v1 Announce Type: new
Abstract: Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the...
By Wenxuan Shen, Dongna Jin, Dongping Chen
arXiv:2608. 11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments.
By Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
SWE-Game is a benchmark comprising 247 tasks based on 41 Godot games across 13 gameplay categories, testing coding agents on tasks such as brief-to-game, design-document implementation, skeleton completion, fault repair, and Godot-to-Unity porting. Evaluation uses engine-state checks, replay of certified reference inputs, and agent-authored demonstrations to judge mechanic correctness, playability, and post-repair behavior, supplemented by vision‑language rubrics for presentation. Across six models, Opus5 leads but overall scores stay below 60/100, highlighting common issues like omitted requirements and gameplay logic errors.
By Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang