arXiv AI

Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection

The paper presents a solution for the UCF UrbanTwin LUMPI Track in the Sim-to-Real LiDAR Challenge, where a detector trained solely on synthetic data must perform on real LiDAR frames. The approach tackles the Sim2Real gap through data alignment, diversified sampling, augmentation, and specialized detectors, followed by class-aware fusion and calibration techniques. The final submission achieved a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and a 3D mAP@0.5 of 0.1258.

arXiv AI
Sep 7

SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

SimFuse3D tackles cross‑platform LiDAR unsupervised domain adaptation by addressing box‑point inconsistency through source‑guided target simulation and confidence‑guided multi‑stage localization reweighting. It preserves target placement, repairs pseudo‑objects using labeled source geometry, and reweights predictions based on confidence, all during adaptation without altering the detector architecture. The method outperforms existing adaptation techniques across six cross‑platform transfers and ranks first on nuScenes‑to‑KITTI for both evaluated detectors.

By Yongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao, Liang Lei, Jianya Guo, Yuqiang Zhai, Xiaofeng Wang, HaiKuo Xu, Haoang Li
arXiv AI
Aug 13

LiDAR-based 3D Change Detection at City Scale

arXiv:2510. 21112v3 Announce Type: replace-cross Abstract: High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and asset monitoring, including both built structures and urban greenery.

By Hezam Albaqami, Haitian Wang, Xinyu Wang, Muhammad Ibrahim, Zainy M. Malakan, Abdullah M. Algamdi, Mohammed H. Alghamdi, Ajmal Mian
arXiv Computer Vision
4d ago

DRS-VPT: Directly Relocalizing in a Scan with Vision Point Transformers

DRS‑VPT is a feed‑forward transformer that performs image‑to‑scan registration by predicting the scan pose, point maps, and a coarse‑to‑fine feature pyramid for direct reprojective alignment. It unifies tasks like camera‑LiDAR calibration and indoor camera‑to‑map relocalization, achieving state‑of‑the‑art results in autonomous driving and competitive performance indoors without map‑specific training. The model also learns complex scan‑to‑image projection behaviors, such as occlusion of back‑facing points.

By Lanke Frank Tarimo Fu, Maurice Fallon
arXiv Computer Vision
Sep 3

GEM: Generating LiDAR World Model via Deformable Mamba

GEM is a Generative LiDAR world model that uses a deformable Mamba architecture to better handle the disorder of LiDAR point clouds and distinguish dynamic objects from static structures. The model tokenizes LiDAR sweeps, unsupervisedly disentangles dynamic and static features, and applies a tri‑path deformable Mamba for selective scanning and adaptive gating fusion, improving spatial‑temporal understanding. Experiments show GEM outperforms existing methods across multiple benchmarks, and it can be paired with a planner and BEV controller for autonomous rollout and "what‑if" scenario generation.

By Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie