arXiv Computer Vision By Hamidreza Mazandarani, Masoud Shokrnezhad, Tarik Taleb, Onur G\"unl\"u

Exploiting Overlapping Fields of View for Redundancy-Aware Uplink Transmission in Vehicular 6G

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv AI
Aug 19

HMS-SCP: Task-Oriented Multi-Scale Semantic Communication for V2X Cooperative Perception

The paper introduces HMS‑SCP, a hierarchical multi‑scale semantic‑aware cooperative perception framework for V2X communication. It uses a spatial importance predictor to select task‑relevant grid elements at multiple scales and maps them directly into complex‑valued symbols for joint source‑channel coding, achieving ultra‑low symbol rates and noise resilience. Experiments on OPV2V and DAIR‑V2X show that HMS‑SCP maintains high‑confidence far‑field detection with sub‑16 ms latency even under severe Rayleigh fading and extreme compression.

By Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
arXiv Machine Learning
Sep 18

QoS-Aware Federated Learning for Multimodal In-Cabin Interaction in Smart Vehicles

The paper introduces FedQoS, an asynchronous federated learning framework designed for multimodal in‑cabin interaction in smart vehicles. It uses a two‑phase gating mechanism: a resource‑aware training gate that starts local learning only when sensing buffers and energy reserves meet safety thresholds, and a QoS‑aware transmission policy that gates uplink updates based on an efficiency score balancing model novelty, latency, and energy costs. Experiments on vehicular datasets show FedQoS achieves competitive personalized accuracy with only marginal loss compared to FedAvg, while reducing communication overhead by 76.7% and latency cost by 26.0%.

By Baran Can G\"ul, Mert Nak{\i}p, Nasser Jazdi, Michael Weyrich
arXiv AI
4d ago

Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs

The paper introduces Encoder‑Sharing Hierarchical Multi‑Task Federated Learning (EN‑HMTFL) for vehicular ad hoc networks, combining cluster‑based hierarchical federated learning with a globally shared encoder and vehicle‑local decoders. This design allows vehicles performing different perception tasks to collaboratively learn a transferable feature representation while keeping raw data and task‑specific decoders local. Experiments on MNIST and GTSRB datasets show that EN‑HMTFL can improve accuracy by up to 24.0% and reduce communication rounds by up to 69 (28.8%) compared to a representation‑sharing benchmark.

By M. Saeid HaghighiFard, Sinem Coleri
arXiv Computer Vision
Sep 21

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.

By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao