arXiv AI

Learned Non-Maximum Suppression for 3D Object Detection

arXiv:2606. 03568v1 Announce Type: cross Abstract: Post-processing is a critical stage in LiDAR-based 3D object detection, where dense and overlapping proposals must be filtered for compact and reliable perception.

arXiv Computer Vision
Sep 18

Open-vocabulary 3D object detection with promptable segmentation

The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.

By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
arXiv AI
Sep 25

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.

By Yuting Zhao, Ziyi Zheng, Shuxiao Li
arXiv AI
Sep 7

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.

By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
arXiv AI
Sep 7

SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

SimFuse3D tackles cross‑platform LiDAR unsupervised domain adaptation by addressing box‑point inconsistency through source‑guided target simulation and confidence‑guided multi‑stage localization reweighting. It preserves target placement, repairs pseudo‑objects using labeled source geometry, and reweights predictions based on confidence, all during adaptation without altering the detector architecture. The method outperforms existing adaptation techniques across six cross‑platform transfers and ranks first on nuScenes‑to‑KITTI for both evaluated detectors.

By Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao, Liang Lei, Jianya Guo, Yuqiang Zhai, Xiaofeng Wang, HaiKuo Xu, Haoang Li