Hugging Face Blog

TextQuests: How Good are LLMs at Text-Based Video Games?

arXiv AI
Sep 28

Game Arena: Strategic LLM Evaluation in Competitive Environments

Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.

By Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen
arXiv Machine Learning
Sep 17

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

Zing-0.5 is a 5B autoregressive world model that enables users to explore and influence generated worlds through joint keyboard and online text control. It integrates unified action and text conditioning, event-scale supervision for incremental generation, and low-cost real-time interaction, achieving high scores on WBench Navigation. The authors release model weights, inference code, and a serving implementation to support further research on playable generated worlds.

By Mingyang Chen, Shengdong Chen, Xiaoxiao Fu, Bosheng Gong, Haoyuan Guo, Bowen Li, Jiawen Li, Kejun Li, Tianpeng Li, Yin Liu, Haoze Sun, Zeyang Tian, Meng Wang, Xinmiao Wu, Jiangqiao Yan, Zining Zhao
arXiv AI
4d ago

WaLLM -- Understanding Use and Engagement with a General-Purpose LLM on WhatsApp

The paper introduces WaLLM, a general‑purpose large language model chatbot deployed on WhatsApp to investigate how users interact with open‑ended AI in everyday settings. The study finds that health and well‑being queries dominate user interactions, and that proactive communication and communal lists—adapted to WhatsApp’s affordances—enhance engagement and content discovery. These insights inform the design of general‑purpose LLM services on messaging platforms.

By Hiba Eltigani, Rukhshan Haroon, Asli Kocak, Abdullah Bin Faisal, Noah Martin, Fahad Dogar