arXiv Computer Vision

UpDown-SC: Gravity-Canonicalized Dual-Envelope Scan Context for Indoor LiDAR Place Recognition

UpDown‑SC is a training‑free polar descriptor for indoor LiDAR place recognition that first canonicalizes gravity and then represents two complementary surfaces: the upper envelope of lower/middle structures and the lower envelope of overhead structures. By estimating a physical split from a cell‑balanced map height distribution and using a mask‑aware, non‑uniform two‑channel distance, it retains discriminative lower‑level evidence while limiting sensitivity to cross‑session variation. Experiments on repeated indoor sessions, mounting‑height changes, mixed outdoor‑to‑indoor trajectories, and an outdoor transfer sequence demonstrate more reliable first‑choice retrieval and significant gains over conventional Scan Context, while maintaining a lightweight CPU front end and supporting metric prior‑map localization.

Hugging Face Trending Papers
Sep 17

SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors (22 089 m²). The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors, and supports three realistic sensing protocols: single walking frames, stationary multi‑view sweeps, and walking streams with odometry. Evaluation of five geometric and learned systems shows that stock models perform poorly, but fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current localization methods.

arXiv Computer Vision
Sep 18

SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors. The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors for trajectory alignment. Five representative systems were evaluated, showing that stock checkpoints perform poorly while fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current methods.

By Yunqian Cheng, Roberto Manduchi
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.