A Survey on Large Language Model-Based Game Agents
arXiv:2404. 02039v5 Announce Type: replace Abstract: Game environments provide rich, controllable settings that stimulate many aspects of real-world complexity.
arXiv:2607. 26393v1 Announce Type: new Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents.
arXiv:2404. 02039v5 Announce Type: replace Abstract: Game environments provide rich, controllable settings that stimulate many aspects of real-world complexity.
arXiv:2604. 11741v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information.
AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.
SocialReasonBench is a new video‑multiple‑choice QA benchmark designed to test socially grounded reasoning in interactive narrative videos. It uses branching gameplay footage from *Detroit: Become Human*, where player choices create alternative social outcomes that can be verified against the game’s script and flowchart. The benchmark includes seven reasoning dimensions—such as intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent—and employs a multi‑agent pipeline to curate clips, ground answer labels, and generate theory‑guided questions with diagnostic distractors.
We present LingBot-World 2. 0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades.
arXiv:2505. 23399v2 Announce Type: replace Abstract: We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning.
arXiv:2509. 00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited information.
arXiv:2605.27068v2 Announce Type: replace-cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Mo...
The paper introduces LM Fight Arena, a new benchmark that pits large multimodal models against each other in the fighting game Mortal Kombat II to evaluate real‑time visual understanding and sequential decision‑making. Six leading open‑ and closed‑source models were tested in a controlled tournament where each controlled the same character, ensuring a fair comparison. The framework offers a fully automated, reproducible, and objective assessment of an LMM’s strategic reasoning in a dynamic setting.
arXiv:2507. 07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously.
The paper introduces VISTA, a visual harness that equips a general-purpose multimodal model with long‑horizon vision and a lossless visual memory. VISTA enables the model to directly perceive and actively retrieve past observations, allowing it to reorganize visual input during reasoning. On the ARC‑AGI‑3 benchmark, VISTA boosts Claude Opus 5.0’s Relative Human Action Efficiency from 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first‑time human participants, and it also outperforms baselines on three additional visual game and puzzle benchmarks.
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp...