arXiv:2503.13938v3 Announce Type: replace-cross
Abstract: Comprehensive traffic scene understanding is a foundational capability for Intelligent Transportation Systems (ITS) underpinning applications...
By Qingyao Xu, Ya Zhang, Yanfeng Wang, Siheng Chen
arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.
By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita
arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.
By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes
arXiv:2608.21440v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail...
By Ran Chen, Jiaxing Ren, Zhikun Zhang, Yunhao Hou, Junbao Zhuo, Bochao Zou
NeuroSymbEAD is a large‑scale neuro‑symbolic caption dataset that builds an ego‑centric knowledge graph of static and dynamic objects on the KITTI‑360 dataset, annotating classes, categories, heading directions, orientations, and distances from the ego‑vehicle. The dataset generates multilevel textual captions that serve as a lightweight representation of an ego‑centric scene map, enabling outdoor scene‑map reconstruction, visual recognition, and object grounding. Baselines for driving common sense and traffic/scene understanding are established, and the dataset is benchmarked using pre‑trained grounding and learned auto‑regressive captioning networks to support vision‑language and foundation models for traffic‑scene explanation, 3D reasoning, and interpretable autonomous‑driving perception.
By Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexplored. In this work, we study the decoding of pedestrian crossing intentions from short egocentric video clips.
arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.
By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
arXiv:2511.20022v3 Announce Type: replace-cross
Abstract: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their...
By Seungjun Yu, Seonho Lee, Namho Kim, Jaeyo Shin, Junsung Park, Wonjeong Ryu, Raehyuk Jung, Hyunjung Shim
arXiv:2608. 16480v1 Announce Type: cross Abstract: We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences.
By Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang
arXiv:2606. 09142v1 Announce Type: cross Abstract: Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexplored.
By Danya Li, Xiang Su, Yan Feng, Rico Krueger
arXiv:2608.28762v1 Announce Type: new
Abstract: Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic sc...
By Shaozu Ding, Linan Song, Dajiang Suo
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
By Chen Yang, Yuhao Wei, Ze Xu, Ziheng Zou, Shuang Liang, Delin Ouyang, Lingfeng Qi, Jie Li, Guofa Li