arXiv AI By Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Read the original on arXiv AI →

Learn2Play Bench is a new benchmark that tests how well large language model agents learn from experience in unfamiliar, text‑based games with novel or counterintuitive rules. The benchmark provides reproducible feedback, automatic scoring, and varied game instances to evaluate learning across repeated attempts and transfer to new situations. Findings show that retaining full action records aids learning, human players outperform agents, and the choice of harness significantly impacts performance and inference cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Oct 1

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Rep2Skill introduces a representation-guided framework that enables large language model agents to self-evolve their textual skills by analyzing internal representation trajectories from agent rollouts. The method identifies execution turns that deviate from successful dynamics and uses these signals, together with execution contexts, as actionable feedback for targeted skill revision. Experiments with two open-source LLMs across two agent environments demonstrate that Rep2Skill consistently outperforms purely text-based approaches, showing that incorporating internal representations can enhance agent self-improvement.

By Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
arXiv AI
1d ago

SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning

SpeedrunBench is a new benchmark that tests large language model agents on their ability to develop and refine strategies for completing nine different video games as quickly as possible. The benchmark requires agents to repeatedly improve, reflect, and reason over long action horizons to beat both themselves and other players. Experiments show that while agents can approach human records on simple platformers, they lag behind humans on longer, more complex games within realistic resource limits.

By Yoshinari Fujinuma, Keisuke Kamahori, Ryuto Koike, Abdelrahman Madkour, Varun Prashant Gangal, Monty Bichouna, Martyna Markiewicz, Shivani Jain, Duncan Curtis, Rebecca Qian, Anand Kannappan
arXiv AI
Jun 30

Hierarchical Experimentalist Agents

arXiv:2606. 29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search.

By Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum