arXiv AI

Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection

arXiv:2605. 20301v2 Announce Type: replace-cross Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making.

arXiv Computer Vision
Sep 23

MatchFusion: Explicit-Implicit Instance Matching for Spatio-Temporal Multimodal Autonomous Driving

MatchFusion is a learnable module that performs explicit-implicit instance matching for spatio‑temporal multimodal autonomous driving. It initializes pairwise affinities with geometric similarity and category consistency, then refines associations using instance embeddings to guide a residual aggregation operator for adaptive information exchange. Experiments on nuScenes show that MatchFusion improves perception accuracy, reduces FLOPs by 55.3% and GPU memory usage by 39.3%, and adds only 3.7% of total perception latency.

By Xiaoyu Li, Jiajia Fu, Long Shi, Tianyu Du, Ruihang Li, Xian Wu, Lijun Zhao, Yingtao Zhang, Lining Sun, Ruifeng Li
arXiv AI
Aug 11

ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking

arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.

By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv Computer Vision
Aug 31

Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation

The paper introduces OccLinker, a lightweight plugin for vision‑based occupancy networks that reduces flickering by efficiently merging historical static and motion cues with current features via a dual cross‑attention mechanism. It generates correction components to refine base network predictions and proposes a new temporal consistency metric to quantify flickering. Experiments on two benchmark datasets show that OccLinker improves performance with minimal computational overhead while effectively diminishing flickering artifacts.

By Fengcheng Yu, Haoran Xu, Canming Xia, Ziyang Zong, Guang Tan
arXiv AI
Sep 7

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.

By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin