Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.
By Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen
PTCG-Bench is a new benchmark that uses the Pokémon Trading Card Game to evaluate large language model (LLM) agents on two fronts: their decision‑making within a single complex game environment and their capacity to evolve through accumulated experience. The benchmark includes a modular harness ablation to isolate agent performance from model capability. Experiments show that while LLM agents can achieve non‑trivial gameplay, sustained self‑evolution remains difficult and performance depends on harness design.
By Dongdong Hua, Yifei Sun, Renhong Huang, Feng Gao, Chunping Wang, Yang Yang