GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range.
HeteroPROMPT is a real‑time, privacy‑preserving framework for heterogeneous collaborative perception in autonomous systems. It aligns features from diverse sensors and models into a unified ego‑centric space using modular prompts and lightweight tuning, while keeping encoders and fusion stacks frozen. The system employs a metadata‑free autoencoder for modality classification and routing, achieving higher average precision on OPV2V‑H and V2XSet datasets with far fewer trainable parameters.
arXiv:2606. 21165v2 Announce Type: replace-cross Abstract: We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving.
arXiv:2609.17856v1 Announce Type: new Abstract: Heterogeneous cooperative perception (CP) enables connected vehicles with diverse sensor setups to share spatial awareness via compact feature maps, wh...
BOLT is a lightweight plug‑and‑play module that enables preparation‑free heterogeneous cooperative perception by adapting neighboring features online through ego‑as‑teacher distillation. It requires only ego predictions, no ground‑truth labels, and uses high‑confidence ego features to align cross‑agent feature domains while allowing neighbors to contribute in low‑confidence regions. With just 0.9 M trainable parameters, BOLT boosts AP@50 by up to 32.3 points over unadapted fusion and consistently outperforms ego‑only results on DAIR‑V2X and OPV2V.
CauseCollab is a causal unified and modality‑agnostic network designed to improve collaborative perception across heterogeneous sensor modalities. It disentangles semantic factors from modality‑specific confounders using causal metric learning and employs a context‑guided Unified Converter to maintain cross‑modal semantic consistency. The approach requires only minimal adapter training when adding new modalities and achieves state‑of‑the‑art results on the OPV2V and DAIR‑V2X datasets, especially in scenarios with large modality gaps.