V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 21165v2 Announce Type: replace-cross Abstract: We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving.
arXiv:2609.36934v1 Announce Type: new Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signal...
arXiv:2608. 07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning.
VIPS is a benchmark for vehicle‑to‑infrastructure cooperative autonomous driving that uses pseudo‑simulation to combine vehicle and infrastructure observations, enabling scalable yet realistic evaluation of robustness and error propagation without full simulation. The paper also introduces CoS‑V2X, a cooperative planning framework that employs sparse representations to model vehicle‑infrastructure interactions efficiently and robustly under heterogeneous observations.
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling.
arXiv:2605.10426v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mecha...