arXiv:2608.21833v1 Announce Type: new
Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especial...
By Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng
arXiv:2609.25001v1 Announce Type: new
Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
By Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
arXiv:2610.08621v1 Announce Type: new
Abstract: Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experi...
By Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu
arXiv:2609.09059v1 Announce Type: new
Abstract: While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existi...
By Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie
arXiv:2609.25652v1 Announce Type: new
Abstract: Recent game world models support realistic visual simulation and interactive gameplay based on player inputs. However, they typically learn environment...
By Zijun Lin, Zhiyang Deng, Yuzhe Wu, Bihan Wen, Yeying Jin
A2Z GameSpec-Bench introduces a benchmark of 100 long‑form Game Design Documents (GDDs) to evaluate how faithfully coding agents can generate complete games from detailed specifications. The benchmark measures faithfulness by checking that the game satisfies the GDD requirements and preserves the relationships among them, using a dependency‑aware contract and a combination of source‑code inspection and agent‑generated test policies. Evaluations show that current agents struggle to meet interdependent requirements, but requirement‑specific feedback improves GDD fidelity by 10.9% after two revision rounds.
By Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park