arXiv AI

A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors

arXiv:2603. 21048v2 Announce Type: replace-cross Abstract: The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the detection of traffic violations and unsafe driver actions.

arXiv Machine Learning
Jun 4

An Open-Source Two-Stage Computer Vision Pipeline for Fine-Grained Vehicle Classification using Vision Transformers

arXiv:2606. 05149v1 Announce Type: cross Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injury-risk-relevant categories from naturalistic roadway video do not exist in the open literature.

By Gandhimathi Padmanaban, Fred Feng
arXiv AI
6d ago

VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

VLALight is a lightweight end‑to‑end vision‑language‑action framework designed for traffic signal control. It fuses multiple camera views and textual instructions to directly predict signal actions, avoiding intermediate image‑to‑text conversions. The model, with only 0.5 B parameters, achieves superior emergency vehicle service, cutting pooled waiting time by 21.1% compared to cascaded methods while running in real time on local hardware.

By Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu, Sicheng Wang, Xi Chen, Yirong Chen, Zhiyong Cui
arXiv AI
Jul 28

From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

arXiv:2510. 03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments.

By Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang