arXiv AI

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V

arXiv AI
Sep 3

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench is an open‑source benchmark that evaluates language‑model agents in the long‑horizon, tool‑mediated game Civilization VI using the Model Context Protocol (MCP). Each episode lasts over 300 turns, generating thousands of tool calls across a 76‑tool action space, and includes a narration layer that translates visual game state into structured text. The study characterises agent behaviour across four model families, introducing Proactive Monitoring Rate (PMR) and RAG@10 as interface‑level metrics, and finds that agents often under‑monitor strategic state and fail to execute near‑term commitments despite tool access and explicit guidance.

By Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa
arXiv AI
6d ago

Game Arena: Strategic LLM Evaluation in Competitive Environments

Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.

By Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen
arXiv AI
Jul 17

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning

arXiv:2606. 29932v3 Announce Type: replace Abstract: Long-horizon strategic planning in complex strategy games requires coordinating tightly coupled decision domains, including technology, economy, diplomacy, and military, across hundreds of turns under imperfect information.

By Tianyu Jin, Shuo Chen, Yida Wang, Liuyu Xiang, Yingzhuo Liu, Zhiyao Jiang, Yexin Li, Peipei Li, Zhaofeng He
arXiv AI
Sep 17

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.

By Joey Xiao, Haonan Huang
arXiv Machine Learning
Sep 4

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

The paper explores a runtime strategy-selection framework where a large language model (LLM) guides a pre‑trained reinforcement learning (RL) policy for non‑player characters (NPCs) in a Unity combat game without altering the underlying policy. Five NPC agents sharing a PPO policy were compared in a baseline setup and an LLM‑augmented setup, where a locally hosted Mistral 7B model assigns one of four tactical tags every five seconds based on live game state. Across 600 episodes against three scripted opponents, the LLM‑augmented agents more than doubled their win rate against a Balanced opponent, improved performance against an Evasive opponent, but struggled against an Aggressive opponent due to over‑reliance on encirclement; analysis of 2,430 strategy selections revealed limited zero‑shot differentiation with the model favoring Surround in 83.8% of cases.

By Hrithika Deepu Nair, Kayvan Karim
arXiv AI
Jun 2

MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution

arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.

By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov
arXiv AI
Sep 17

LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

The paper introduces LM Fight Arena, a new benchmark that pits large multimodal models against each other in the fighting game Mortal Kombat II to evaluate real‑time visual understanding and sequential decision‑making. Six leading open‑ and closed‑source models were tested in a controlled tournament where each controlled the same character, ensuring a fair comparison. The framework offers a fully automated, reproducible, and objective assessment of an LMM’s strategic reasoning in a dynamic setting.

By Yushuo Zheng, Tongrui Ye, Zicheng Zhang, Xiongkuo Min, Huiyu Duan, Guangtao Zhai