arXiv AI

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

arXiv:2608. 08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling.

arXiv Computation and Language
Sep 7

ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

ConWriter is a training‑free framework that generates long‑form stories scene by scene, using static requirements, dynamic memory, symbolic state reasoning, and uncertainty‑aware risk signals to enforce consistency. It checks each new scene against required narrative transitions and repairs local errors before they propagate. Evaluations on ConStory‑Bench show that ConWriter matches or outperforms direct generation and a recent baseline, improving narrative consistency across multiple models and story lengths.

By Jindong Li, Yang Yang, Zihao Liu, Yutao Yue, Menglin Yang
arXiv AI
Jun 16

Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP

arXiv:2606. 16014v1 Announce Type: cross Abstract: Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated worlds -- most acute in sandbox and open-world settings -- has been prohibitively expensive.

By Yuhang Huang, Chenmiao Li, Chaowei Fang
arXiv AI
Sep 3

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.

By Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa