SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors. The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors for trajectory alignment. Five representative systems were evaluated, showing that stock checkpoints perform poorly while fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current methods.
By Yunqian Cheng, Roberto Manduchi
UpDown‑SC is a training‑free polar descriptor for indoor LiDAR place recognition that first canonicalizes gravity and then represents two complementary surfaces: the upper envelope of lower/middle structures and the lower envelope of overhead structures. By estimating a physical split from a cell‑balanced map height distribution and using a mask‑aware, non‑uniform two‑channel distance, it retains discriminative lower‑level evidence while limiting sensitivity to cross‑session variation. Experiments on repeated indoor sessions, mounting‑height changes, mixed outdoor‑to‑indoor trajectories, and an outdoor transfer sequence demonstrate more reliable first‑choice retrieval and significant gains over conventional Scan Context, while maintaining a lightweight CPU front end and supporting metric prior‑map localization.
By Jie Xu, Yongxin Yang, Ziyi Jin, Kangjin Yu, Hongjun Huang, Chao Han, Zhongpu Xia
arXiv:2608.29426v1 Announce Type: cross
Abstract: Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonom...
By Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Fr\'ed\'eric Kaplan
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
By Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
arXiv:2608.30657v1 Announce Type: new
Abstract: Fixed-viewpoint infrastructure sensors repeatedly observe the same traffic space, making roadside 3D occupancy structurally different from ego-vehicle...
By Lei Yang, Xiaokai Bai, Boqi Li, Chunmian Lin, Li Wang, Ziying Song, Jiahuan Zhang, Enhui Ma, Haibao Yu, Jiaqi Ma, Kaicheng Yu
The paper investigates whether progress on spatial reasoning benchmarks translates into improved navigation performance. It finds a gap between benchmark-oriented spatial specialization and real navigation, and proposes aligning spatial supervision with navigation goals, phases, and decision learning. The authors introduce Spatial-Nav-100K, a two-stage fine-tuning approach, and Spatial-NPD, a teacher–student framework that eliminates explicit spatial reasoning at inference, achieving state‑of‑the‑art success rates on HM3D and MP3D with efficient GPU usage.
By Xun Huang, Shijia Zhao, Rongsheng Qu, Jiayuan Li, Xin Lu, Weixin Li, Chenglu Wen, Cheng Wang
arXiv:2603. 04277v2 Announce Type: replace-cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene.
By Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tr...
arXiv:2603. 16016v2 Announce Type: replace-cross Abstract: A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surroundings would better serve applications such as indoor navigation.
By Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome
VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.
By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv:2607.04745v2 Announce Type: replace-cross
Abstract: Strong mapped-region thermal visual place recognition (VPR) does not ensure safe rejection of unmapped queries. We identify and quantify this...
By Zhiyuan Lu, Kanji Tanaka
arXiv:2609.34302v2 Announce Type: replace
Abstract: Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often req...
By Taigo Sakai, Hiroki Kouno, Naoki Kato, Kazuhiro Hotta