arXiv Computer Vision

Route-MHT: Multimodal Transformer Guardrails for Thermal Visual Place Recognition

arXiv Computer Vision
Sep 23

Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

The paper introduces TFA, a training‑free aggregation technique that calibrates frozen visual foundation models for visual place recognition. TFA uses cross‑codebook agreement, retrieval coverage, and spectral statistics to adjust residual assignment, spectral shaping, and global‑feature fusion without requiring place labels or task‑specific weights. Experiments with a DINOv2‑B backbone show significant Recall@1 gains over existing training‑free methods across multiple benchmarks, demonstrating that reliability‑guided aggregation can unlock additional retrieval performance from frozen representations.

By Xin Li, Zhimin Mao, Shang Wang, Siyuan Duan, Geng Zhang
Hugging Face Trending Papers
Jul 2

DL-VINS-Factory: A Modular Framework for Learned Visual Front-Ends in Visual-Inertial SLAM

Deep-learning features excel in visual matching, yet their practical value in tightly coupled visual-inertial SLAM (VI-SLAM) remains insufficiently characterized. We present DL-VINS-Factory, a unified framework that integrates learned feature extractors (ALIKED, RaCo, SuperPoint, XFeat) with either Lucas--Kanade (LK) optical-flow tracking or LightGlue (LG) descriptor matching.

arXiv AI
Aug 21

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

arXiv:2608. 19536v1 Announce Type: cross Abstract: Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics.

By Eunsoo Im, Junghun Suh, Gyeonggwan Lee, Seunghwan Hong
arXiv Computer Vision
Sep 21

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

By Adam D. Hines, Michael Milford, Tobias Fischer
Hugging Face Trending Papers
Sep 17

SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors (22 089 m²). The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors, and supports three realistic sensing protocols: single walking frames, stationary multi‑view sweeps, and walking streams with odometry. Evaluation of five geometric and learned systems shows that stock models perform poorly, but fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current localization methods.

arXiv Computer Vision
Sep 24

Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM

Know-Your-Scene (KYS)-SLAM extends ORB‑SLAM3 by replacing binary feature rejection with continuous correspondence modulation based on semantic, panoptic, and motion priors. Each keypoint is augmented with hierarchical compatibility scores that down‑weight features on independently moving objects while preserving static structure, using a training‑free depth‑aware ego‑motion model and self‑calibrating thresholds. Across 21 stereo sequences, KYS‑SLAM achieves a 17.4% ATE RMSE reduction on outdoor KITTI, 27.7% on indoor EuRoC, and significant improvements on dynamic and synthetic datasets without per‑sequence tuning.

By Preeti Chatterjee, Jin Lu, Jin Sun, Suchendra M. Bhandarkar
arXiv Computer Vision
Sep 21

Refining Ground Truth Poses in Autonomous Driving Datasets via Neural Rendering

arXiv:2504.15776v2 Announce Type: replace Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...

By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv Computer Vision
1d ago

UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

arXiv:2610.00878v1 Announce Type: cross Abstract: General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary ini...

By Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang