arXiv AIBy Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen
Game Arena: Strategic LLM Evaluation in Competitive Environments
Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.
Machine-generated by The Flow from the publisher's headline and feed description
— not written or checked by a human. The full article lives at arXiv AI.
The paper introduces LM Fight Arena, a new benchmark that pits large multimodal models against each other in the fighting game Mortal Kombat II to evaluate real‑time visual understanding and sequential decision‑making. Six leading open‑ and closed‑source models were tested in a controlled tournament where each controlled the same character, ensuring a fair comparison. The framework offers a fully automated, reproducible, and objective assessment of an LMM’s strategic reasoning in a dynamic setting.
By Yushuo Zheng, Tongrui Ye, Zicheng Zhang, Xiongkuo Min, Huiyu Duan, Guangtao Zhai
arXiv:2609.25001v1 Announce Type: new
Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
arXiv:2608. 04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike.
By S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath
arXiv:2607. 24573v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult.
By Jonas Schr\"oder, Jonas Schweisthal, Oliver M\"uller, Markus Weinmann, Stefan Feuerriegel