arXiv Computer Vision

GA-EIRFS: A Geometry-Augmented Repeat-Factor Sampling Method for Long-Tailed LiDAR 3D Object Detection

GA-EIRFS is a detector‑agnostic sampling strategy that augments frequency‑based repeat‑factor sampling with a fixed geometry score derived from point count, surface‑normal entropy, and surface coverage. It only alters frame‑sampling probabilities, leaving the underlying detector and inference pipeline unchanged. Experiments on nuScenes show consistent improvements in mean average precision and the nuScenes detection score across multiple seeds and backbones, with notable gains for rare classes such as bicycles.

arXiv Computer Vision
Sep 18

Open-vocabulary 3D object detection with promptable segmentation

The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.

By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
arXiv AI
Sep 7

SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

SimFuse3D tackles cross‑platform LiDAR unsupervised domain adaptation by addressing box‑point inconsistency through source‑guided target simulation and confidence‑guided multi‑stage localization reweighting. It preserves target placement, repairs pseudo‑objects using labeled source geometry, and reweights predictions based on confidence, all during adaptation without altering the detector architecture. The method outperforms existing adaptation techniques across six cross‑platform transfers and ranks first on nuScenes‑to‑KITTI for both evaluated detectors.

By Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao, Liang Lei, Jianya Guo, Yuqiang Zhai, Xiaofeng Wang, HaiKuo Xu, Haoang Li
arXiv AI
Sep 10

Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection

The paper presents a solution for the UCF UrbanTwin LUMPI Track in the Sim-to-Real LiDAR Challenge, where a detector trained solely on synthetic data must perform on real LiDAR frames. The approach tackles the Sim2Real gap through data alignment, diversified sampling, augmentation, and specialized detectors, followed by class-aware fusion and calibration techniques. The final submission achieved a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and a 3D mAP@0.5 of 0.1258.

By Pu Luo, Cong Xu, Yumei Li, Kexin Zhang, Licheng Jiao, Wenping Ma, Lingling Li
arXiv Computer Vision
Sep 25

UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition

UpDown‑SC is a training‑free polar descriptor for indoor LiDAR place recognition that first canonicalizes gravity and then represents two complementary surfaces: the upper envelope of lower/middle structures and the lower envelope of overhead structures. By estimating a physical split from a cell‑balanced map height distribution and using a mask‑aware, non‑uniform two‑channel distance, it retains discriminative lower‑level evidence while limiting sensitivity to cross‑session variation. Experiments on repeated indoor sessions, mounting‑height changes, mixed outdoor‑to‑indoor trajectories, and an outdoor transfer sequence demonstrate more reliable first‑choice retrieval and significant gains over conventional Scan Context, while maintaining a lightweight CPU front end and supporting metric prior‑map localization.

By Jie Xu, Yongxin Yang, Ziyi Jin, Kangjin Yu, Hongjun Huang, Chao Han, Zhongpu Xia