arXiv:2608. 16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images.
By Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted to low-level vision schemes, struggling to perceive scene semantics and compositions, which limits both interpretability and applicability.
arXiv:2609.09396v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...
By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
arXiv:2512. 02473v2 Announce Type: replace-cross Abstract: Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions.
By Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
Know-Your-Scene (KYS)-SLAM extends ORB‑SLAM3 by replacing binary feature rejection with continuous correspondence modulation based on semantic, panoptic, and motion priors. Each keypoint is augmented with hierarchical compatibility scores that down‑weight features on independently moving objects while preserving static structure, using a training‑free depth‑aware ego‑motion model and self‑calibrating thresholds. Across 21 stereo sequences, KYS‑SLAM achieves a 17.4% ATE RMSE reduction on outdoor KITTI, 27.7% on indoor EuRoC, and significant improvements on dynamic and synthetic datasets without per‑sequence tuning.
By Preeti Chatterjee, Jin Lu, Jin Sun, Suchendra M. Bhandarkar
The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.
By Adam D. Hines, Michael Milford, Tobias Fischer