Training a Conditioned Video Game Agent on a VLM Annotated Dataset
arXiv:2608. 05954v1 Announce Type: new Abstract: Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning.
arXiv:2608. 05949v1 Announce Type: new Abstract: Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications.
arXiv:2608. 05954v1 Announce Type: new Abstract: Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning.
arXiv:2602. 13602v2 Announce Type: replace-cross Abstract: We present \revise (\underline{Re}asoning with \underline{Vi}deo \underline{S}parsity), a multi-round agent for video question answering (VQA).
arXiv:2607. 25921v1 Announce Type: cross Abstract: In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping.
The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.
arXiv:2606. 20210v1 Announce Type: new Abstract: Immersion in video games depends not only on graphics, audio, and game mechanics, but also on the quality of in-game characters.
arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.
arXiv:2607. 02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process.
In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels.
arXiv:2512. 05277v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.
arXiv:2512. 05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.
arXiv:2609.25001v1 Announce Type: new Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
The paper introduces Instruct-to-Act, a system that decouples high‑level planning from low‑latency control by combining a vision‑language model (VLM) planner with a world‑model controller. The VLM generates sparse, high‑level text instructions, while the controller executes them autonomously at high frequency. Experiments across seven embodied environments, including multi‑agent settings, show that this approach outperforms both controller‑only and direct VLM action‑generation methods, maintains fast control, and allows swapping in different pretrained VLM planners without fine‑tuning.