arXiv AI

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

arXiv:2608. 07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning.

arXiv Computer Vision
3d ago

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

arXiv:2609.38641v1 Announce Type: new Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...

By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
arXiv Computer Vision
Sep 3

VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation

VIPS is a benchmark for vehicle‑to‑infrastructure cooperative autonomous driving that uses pseudo‑simulation to combine vehicle and infrastructure observations, enabling scalable yet realistic evaluation of robustness and error propagation without full simulation. The paper also introduces CoS‑V2X, a cooperative planning framework that employs sparse representations to model vehicle‑infrastructure interactions efficiently and robustly under heterogeneous observations.

By Hoonhee Cho, Jae-Young Kang, Giwon Lee, Hyemin Yang, Heejun Park, Kuk-Jin Yoon
arXiv Computer Vision
Sep 21

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.

By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
arXiv AI
Jun 4

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

arXiv:2512. 05277v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.

By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari
arXiv AI
Jun 2

From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model

arXiv:2512. 05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of the most safety-critical instances.

By Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain, Ahmad Rezaei, Mohsen Gholami, Alireza Heidarikhazaei, Zhou Weimin, Yong Zhang, Mohammad Akbari