Learn2Play Bench is a new benchmark that tests how well large language model agents learn from experience in unfamiliar, text‑based games with novel or counterintuitive rules. The benchmark provides reproducible feedback, automatic scoring, and varied game instances to evaluate learning across repeated attempts and transfer to new situations. Findings show that retaining full action records aids learning, human players outperform agents, and the choice of harness significantly impacts performance and inference cost.
By Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi
arXiv:2608. 09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents.
By Keyu He, Xuhui Zhou, Maarten Sap
arXiv:2510. 11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play.
By Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv:2608. 11338v1 Announce Type: cross Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence.
By Zixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jim\'enez Guti\'errez, Daniel Khashabi, Nicholas Andrews
arXiv:2604. 02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI.
By DeepReinforce Team, Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li
UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.
By Wenjie Liao, Liangjie Zhao, Zehong Cao