arXiv AI

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

arXiv:2510. 23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors.

arXiv AI
Jul 2

Coachable agents for interactive gameplay

arXiv:2607. 00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models.

By Roberto Capobianco (Sony AI, Zurich, Switzerland), Harm van Seijen (Sony AI, North America, various locations), Nolan D. Bard (Sony AI, North America, various locations), Neil Burch (Sony AI, North America, various locations), Fatima Davelouis (Sony AI, North America, various locations), Josh Davidson (Sony AI, North America, various locations), Alisa Devlic (Sony AI, Zurich, Switzerland), Yunshu Du (Sony AI, North America, various locations), Ishan Durugkar (Sony AI, North America, various locations), Siddhant Gangapurwala (Sony AI, North America, various locations), Daniel Hernandez (Sony AI, North America, various locations), G. Zacharias Holland (Sony AI, North America, various locations), Sahil Jain (Sony AI, North America, various locations), Kenta Kawamoto (Sony AI, Tokyo, Japan), Raksha Kumaraswamy (Sony AI, North America, various locations), Patrick MacAlpine (Sony AI, North America, various locations), Dustin R. Morrill (Sony AI, North America, various locations), Declan Oller (Sony AI, North America, various locations), Francesco Riccio (Sony AI, Zurich, Switzerland), Akanksha Saran (Sony AI, North America, various locations), Craig Sherstan (Sony AI, Tokyo, Japan), Kaushik Subramanian (Sony AI, Zurich, Switzerland), Thomas J. Walsh (Sony AI, North America, various locations), Samuel Barrett (Sony AI, North America, various locations), Kizza N. Frisbee (Sony AI, North America, various locations), Mady Govil (Sony AI, North America, various locations), Johannes G\"unther (Sony AI, North America, various locations), Varun R. Kompella (Sony AI, North America, various locations), James A. MacGlashan (Sony AI, North America, various locations), Maxwell Svetlik (Sony AI, North America, various locations), Michael D. Thomure (Sony AI, North America, various locations), Jaden B. Travnik (Sony AI, North America, various locations), Kevin Waugh (Sony AI, North America, various locations), Elahe Aghapour (Sony AI, North America, various locations), Florian Fuchs (Sony AI, Zurich, Switzerland), Andreanne Lemay (Sony AI, North America, various locations), Shruti Mishra (Sony AI, Zurich, Switzerland), Takuma Seno (Sony AI, Tokyo, Japan), Peter Stone (Sony AI, North America, various locations), Michael Spranger (Sony AI, Tokyo, Japan), Peter R. Wurman (Sony AI, North America, various locations)
arXiv AI
Sep 17

LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

The paper introduces LM Fight Arena, a new benchmark that pits large multimodal models against each other in the fighting game Mortal Kombat II to evaluate real‑time visual understanding and sequential decision‑making. Six leading open‑ and closed‑source models were tested in a controlled tournament where each controlled the same character, ensuring a fair comparison. The framework offers a fully automated, reproducible, and objective assessment of an LMM’s strategic reasoning in a dynamic setting.

By Yushuo Zheng, Tongrui Ye, Zicheng Zhang, Xiongkuo Min, Huiyu Duan, Guangtao Zhai
arXiv AI
Sep 23

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?

PTCG-Bench is a new benchmark that uses the Pokémon Trading Card Game to evaluate large language model (LLM) agents on two fronts: their decision‑making within a single complex game environment and their capacity to evolve through accumulated experience. The benchmark includes a modular harness ablation to isolate agent performance from model capability. Experiments show that while LLM agents can achieve non‑trivial gameplay, sustained self‑evolution remains difficult and performance depends on harness design.

By Dongdong Hua, Yifei Sun, Renhong Huang, Feng Gao, Chunping Wang, Yang Yang