RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
arXiv:2606. 18950v1 Announce Type: new Abstract: Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.
The paper introduces LM Fight Arena, a new benchmark that pits large multimodal models against each other in the fighting game Mortal Kombat II to evaluate real‑time visual understanding and sequential decision‑making. Six leading open‑ and closed‑source models were tested in a controlled tournament where each controlled the same character, ensuring a fair comparison. The framework offers a fully automated, reproducible, and objective assessment of an LMM’s strategic reasoning in a dynamic setting.
arXiv:2606. 18950v1 Announce Type: new Abstract: Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.
arXiv:2507. 07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously.
Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action semantics depend on coordinated roles and spatial-te...
The paper introduces DoublesEval, a diagnostic framework that uses professional doubles badminton to test visual‑language models’ ability to reason about dynamic multi‑agent interactions. It decomposes rallies into key moments and evaluates models across four dimensions—atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction—highlighting specific reasoning failures. The authors also propose TacticCheck, a lightweight consistency checker that improves performance without retraining the models, yet significant gaps remain in tactical reasoning.
arXiv:2607. 01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance.
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
arXiv:2606. 09826v1 Announce Type: cross Abstract: Vision-language model (VLM) agents are increasingly deployed in interactive game environments.
arXiv:2601. 02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex reasoning.
PAVXploreRL introduces a reinforcement learning framework that builds on a pretrained latent world model to explicitly optimize Physical Plausibility, Action Adherence, and Visual Fidelity (PAV) objectives. By combining in‑distribution expert trajectories with noise‑driven out‑of‑distribution action exploration, the method avoids reliance on paired video supervision and improves generalization. Experiments demonstrate a 5.6% average performance gain over pretrained baselines and more reliable policy evaluation with reduced overestimation bias.
arXiv:2607. 26393v1 Announce Type: new Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents.
arXiv:2505. 23399v2 Announce Type: replace Abstract: We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning.
arXiv:2510. 23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors.