arXiv AI By Yoshinari Fujinuma, Keisuke Kamahori, Ryuto Koike, Abdelrahman Madkour, Varun Prashant Gangal, Monty Bichouna, Martyna Markiewicz, Shivani Jain, Duncan Curtis, Rebecca Qian, Anand Kannappan

SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning

Read the original on arXiv AI →

SpeedrunBench is a new benchmark that tests large language model agents on their ability to develop and refine strategies for completing nine different video games as quickly as possible. The benchmark requires agents to repeatedly improve, reflect, and reason over long action horizons to beat both themselves and other players. Experiments show that while agents can approach human records on simple platformers, they lag behind humans on longer, more complex games within realistic resource limits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Learn2Play Bench is a new benchmark that tests how well large language model agents learn from experience in unfamiliar, text‑based games with novel or counterintuitive rules. The benchmark provides reproducible feedback, automatic scoring, and varied game instances to evaluate learning across repeated attempts and transfer to new situations. Findings show that retaining full action records aids learning, human players outperform agents, and the choice of harness significantly impacts performance and inference cost.

By Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi
arXiv AI
Jul 14

People use fast and flat simulation to reason about new games

arXiv:2510. 11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play.

By Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao