Robust dynamic object detection and tracking are essential for enabling robots to operate safely and effectively alongside humans in complex environments such as construction sites. While LiDAR-based SLAM and occupancy grid methods offer viable solutions for detecting and tracking motion, many state-of-the-art 3D vision approaches rely heavily on pre-trained neural networks and require additional post-processing to identify moving objects.
arXiv:2607.11099v2 Announce Type: replace-cross
Abstract: Reliable visual data association is fundamental to visual SLAM (V-SLAM), as it directly determines the quality of the camera pose estimation...
By Ting-Wei Ou, Huang-Ting Lin, Kuu-Young Young
AMB3R‑SLAM is a real‑time monocular SLAM system that can reconstruct kilometer‑scale trajectories over 10,000 frames on a single consumer‑grade GPU. It combines a lightweight front‑end for low‑latency tracking with a hierarchical backend that enforces local, mid‑level, and global consistency, avoiding bundle adjustment and thus handling dynamic scenes naturally. The system also supports stereo, RGB‑D, and LiDAR inputs, achieving strong camera tracking performance and reducing absolute trajectory error by over 70% on several datasets, with sub‑meter accuracy when LiDAR is added.
By Hengyi Wang, Lourdes Agapito
arXiv:2405.07392v4 Announce Type: replace-cross
Abstract: Many existing visual SLAM methods can achieve high localization accuracy in dynamic environments by leveraging deep learning to mask moving o...
By Yuhao Zhang, Mihai Bujanca, Mikel Luj\'an
arXiv:2607.24495v2 Announce Type: replace
Abstract: Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their d...
By Jiaheng Li, Binsheng Zhang, Xinhai Chang, Wenzheng Chen
The paper presents a framework that builds a static point cloud prior map from past camera traversals, augmenting each point with DINOv3 semantic features. During runtime, a local prior patch is retrieved, encoded with a sparse voxel backbone, and fused with lifted multi‑view camera features in bird’s‑eye view. This fused representation is then used by sparse transformer heads to predict 3D objects and vectorized map elements, achieving improved performance on Argoverse 2 without requiring LiDAR for prior‑map construction or online inference.
By Markus K\"appeler, Rohit Mohan, Abhinav Valada
SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.
By Yuting Zhao, Ziyi Zheng, Shuxiao Li
3D scene graphs provide a hierarchical abstraction of environments by encoding spatial entities, such as objects and places, and their relationships. However, existing scene graph systems model object geometry coarsely, relying on partial point clouds or class-level CAD templates, which limits instance-specific shape detail.
arXiv:2608.13147v2 Announce Type: replace
Abstract: Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera str...
By Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li, Ya Yang, Naiyan Wang, Si Liu
arXiv:2607. 23384v1 Announce Type: cross Abstract: Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy depends critically on associating measurements with the correct landmark variables.
By Yihao Zhang, Jungseok Hong, John J. Leonard
DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
The paper introduces Post Fusion Stabilizer (PFS), a lightweight module that refines intermediate bird’s‑eye view (BEV) feature maps in existing camera‑LiDAR fusion detectors. PFS stabilizes feature statistics under domain shift, suppresses regions affected by sensor degradation, and adaptively restores weakened cues via residual correction, acting as a near‑identity transformation. On the nuScenes benchmark, PFS achieves state‑of‑the‑art robustness, notably improving camera dropout robustness by +1.2% and low‑light performance by +4.4% mAP while adding only 3.3 M parameters.
By Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin