arXiv:2608. 14603v1 Announce Type: cross Abstract: Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots.
By Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
The paper introduces HMS‑SCP, a hierarchical multi‑scale semantic‑aware cooperative perception framework for V2X communication. It uses a spatial importance predictor to select task‑relevant grid elements at multiple scales and maps them directly into complex‑valued symbols for joint source‑channel coding, achieving ultra‑low symbol rates and noise resilience. Experiments on OPV2V and DAIR‑V2X show that HMS‑SCP maintains high‑confidence far‑field detection with sub‑16 ms latency even under severe Rayleigh fading and extreme compression.
By Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.
By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui
arXiv:2609.37098v1 Announce Type: cross
Abstract: Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providin...
By Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, Bin Ran
Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.
arXiv:2608. 04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems.
By Yu Zhao, Jiangyu Pan, Tao Hu, Ming Yin, Fan Yang, Jiangfan Liu, Xiubo Liang
arXiv:2609.36934v1 Announce Type: new
Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signal...
By Pan Zhang, Siqi Lai, Kemu Dong, Hao Liu
arXiv:2503.13938v3 Announce Type: replace-cross
Abstract: Comprehensive traffic scene understanding is a foundational capability for Intelligent Transportation Systems (ITS) underpinning applications...
By Qingyao Xu, Ya Zhang, Yanfeng Wang, Siheng Chen
The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to jointly perform heterogeneous classification and regression tasks on the Cityscapes dataset, using an InfoMax principle to handle mixed discrete and continuous semantic variables. It compares the new framework against independent single‑task training, conventional task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom, and studies how the capacity of the common unit affects joint task performance. Extensive evaluations show that CMT‑SemCom outperforms all benchmarks.
By Ahmad Halimi Razlighi, Mohammad Siddiqur Rahman, Maximilian H. V. Tillmann, Edgar Beck, Armin Dekorsy
arXiv:2606. 15749v1 Announce Type: cross Abstract: Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics.
By Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen
VLALight is a lightweight end‑to‑end vision‑language‑action framework designed for traffic signal control. It fuses multiple camera views and textual instructions to directly predict signal actions, avoiding intermediate image‑to‑text conversions. The model, with only 0.5 B parameters, achieves superior emergency vehicle service, cutting pooled waiting time by 21.1% compared to cascaded methods while running in real time on local hardware.
By Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu, Sicheng Wang, Xi Chen, Yirong Chen, Zhiyong Cui
arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.
By Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian