arXiv AI By Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Aug 28

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.

By Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Gongshen Liu, Zhuosheng Zhang
arXiv AI
Sep 21

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

arXiv:2609.21562v1 Announce Type: cross Abstract: Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A gam...

By Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei, Yanghai Wang, Zixuan Dong, Yifan Yao, Qianqian Xie, Letian Zhu, Jiaheng Liu
arXiv AI
Jun 29

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.

By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
arXiv AI
Sep 3

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.

By Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa