arXiv:2512. 05277v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.
By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv:2608. 19380v1 Announce Type: new Abstract: While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpret traffic accidents.
By Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich
Latent-CoT-Drive (LCDrive) is a vision‑language‑action model for autonomous driving that replaces natural‑language chain‑of‑thought reasoning with a latent language capturing possible outcomes of driving actions. The model interleaves action‑proposal tokens, aligned with the model’s output actions, and world‑model tokens grounded in a learned latent world model to reason about future outcomes. After a supervised cold‑start using ground‑truth future rollouts, LCDrive is further refined with closed‑loop reinforcement learning, achieving faster inference, higher‑quality trajectories, and greater gains from interactive RL than both non‑reasoning and text‑reasoning baselines on a large‑scale end‑to‑end driving benchmark.
By Shuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian, Yurong You, Yan Wang, Wenjie Luo, Yulong Cao, Philipp Krahenbuhl, Marco Pavone, Boris Ivanovic
arXiv:2608. 04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions.
By Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang
arXiv:2609.38641v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...
By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
AgentVidBench is a new multi‑hop video question‑answering benchmark designed to evaluate spatial, temporal, and causal reasoning in multimodal large language models (MLLMs). Unlike existing tests that focus on simple scene queries or global summaries, AgentVidBench includes step‑by‑step solution traces to assess whether agents gather the necessary evidence to justify their answers. Experiments with 12 MLLMs show limited single‑turn performance, but integrating these models into agentic workflows improves both accuracy and trajectory scores, establishing AgentVidBench as a comprehensive testbed for future research on agentic video understanding.
By Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee