arXiv AI

VLALight: A Vision-Language-Action Model for Traffic Signal Control

arXiv AI
6d ago

VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

VLALight is a lightweight end‑to‑end vision‑language‑action framework designed for traffic signal control. It fuses multiple camera views and textual instructions to directly predict signal actions, avoiding intermediate image‑to‑text conversions. The model, with only 0.5 B parameters, achieves superior emergency vehicle service, cutting pooled waiting time by 21.1% compared to cascaded methods while running in real time on local hardware.

By Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu, Sicheng Wang, Xi Chen, Yirong Chen, Zhiyong Cui
arXiv AI
Jun 16

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.

By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
arXiv AI
Jun 2

TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control

arXiv:2604. 17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, yet extending them to physically grounded systems remains challenging.

By Siqi Lai, Pan Zhang, Yuping Zhou, Jindong Han, Yansong Ning, Hao Liu
arXiv AI
Aug 11

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

arXiv:2608. 07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning.

By Hsu-kuang Chiu, Stephen F. Smith
Hugging Face Trending Papers
Aug 18

Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.

arXiv AI
6d ago

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

TrafficImag is the first benchmark designed to evaluate counterfactual roadside traffic video generation, combining a large roadside dataset with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is encoded as an actor-level program specifying target actor, intended behavior, legal route, interaction order, and temporal constraints, allowing a unified evaluation across diverse foundation models. The benchmark assesses four validity dimensions—initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation—and reports that the best models achieve 80.4% macro F1 for reasoning and 55.0% end-to-end success when using a complete condition interface.

By Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
arXiv Computer Vision
4d ago

CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.

By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue