arXiv AI By Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu, Sicheng Wang, Xi Chen, Yirong Chen, Zhiyong Cui

VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

Read the original on arXiv AI →

VLALight is a lightweight end‑to‑end vision‑language‑action framework designed for traffic signal control. It fuses multiple camera views and textual instructions to directly predict signal actions, avoiding intermediate image‑to‑text conversions. The model, with only 0.5 B parameters, achieves superior emergency vehicle service, cutting pooled waiting time by 21.1% compared to cascaded methods while running in real time on local hardware.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.

By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
arXiv Machine Learning
Sep 24

Less Language, More Latents: Annotation-Efficient VLAs for Driving

The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.

By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
arXiv AI
Jul 10

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

arXiv:2607. 08745v1 Announce Type: new Abstract: Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering.

By Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, Jugal Kalita