LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.
By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv:2503.13938v3 Announce Type: replace-cross
Abstract: Comprehensive traffic scene understanding is a foundational capability for Intelligent Transportation Systems (ITS) underpinning applications...
By Qingyao Xu, Ya Zhang, Yanfeng Wang, Siheng Chen
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
arXiv:2602. 23499v4 Announce Type: replace-cross Abstract: Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render the entire dataset unusable.
By Tugrul Gorgulu, Atakan Dag, M. Esat Kalfaoglu, Halil Ibrahim Kuru, Baris Can Cam, Halil Ibrahim Ozturk, Ozsel Kilinc
arXiv:2603. 06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios.
By Thomas Monninger, Shaoyuan Xie, Qi Alfred Chen, Sihao Ding
RiskWorld is a risk‑aware world modeling framework that forecasts shared occupancy and selectively replaces planned trajectories in automated driving. It fuses spatial risk fields, temporal actor context, and visual bird’s‑eye‑view features, using flow‑guided evolution to transport occupancy and signed residuals to correct it. In open‑loop planning on nuScenes, RiskWorld achieves the lowest collision rate over a 3‑second horizon and the second‑best average L2 error, running at 11.5 FPS on a single NVIDIA RTX 4090.
By Rongxiang Zeng, Linsen Cai, Jiafu Zhang, Yijie Zhong, Yide Tao, Shuai Wang, Nan Zheng, Hai L. Vu, Alvaro Garcia Hernandez, Yongqi Dong