DefVINS is a visual‑inertial odometry pipeline tailored for deformable scenes, breaking the rigidity assumption of traditional VIO. It decomposes the odometry state into a rigid, IMU‑anchored component and a non‑rigid scene warp using an embedded deformation graph. The authors also introduce VIMandala, the first real‑world benchmark with ground‑truth camera poses for deformable VIO, and extend the synthetic Drunkard’s benchmark with inertial data, demonstrating that DefVINS outperforms both rigid and non‑rigid baselines.
By Samuel Cerezo, Javier Civera
arXiv:2609.38054v1 Announce Type: cross
Abstract: We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MA...
By Christopher Kolios, Ishaan Mehta, Sasa Janjic, Yeganeh Bahoo, Sajad Saeedi
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv:2607.11099v2 Announce Type: replace-cross
Abstract: Reliable visual data association is fundamental to visual SLAM (V-SLAM), as it directly determines the quality of the camera pose estimation...
By Ting-Wei Ou, Huang-Ting Lin, Kuu-Young Young
arXiv:2609.15795v1 Announce Type: new
Abstract: Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issu...
By Mingkai Liu, Hao Zhao, Xingxing Zuo
arXiv:2609.15795v2 Announce Type: replace
Abstract: Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental...
By Mingkai Liu, Hao Zhao, Xingxing Zuo
arXiv:2509.04600v2 Announce Type: replace
Abstract: Reconstructing global human motion from monocular video is fundamental to VR, graphics, and robotics, yet remains ill-posed due to depth ambiguity,...
By Zhongyuan Hu, Qijun Ying, Jiazhi Shu, Ronghui Li, Yu Lu, Zijiao Zeng, Xiu Li
arXiv:2607.24495v2 Announce Type: replace
Abstract: Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their d...
By Jiaheng Li, Binsheng Zhang, Xinhai Chang, Wenzheng Chen
Deep-learning features excel in visual matching, yet their practical value in tightly coupled visual-inertial SLAM (VI-SLAM) remains insufficiently characterized. We present DL-VINS-Factory, a unified framework that integrates learned feature extractors (ALIKED, RaCo, SuperPoint, XFeat) with either Lucas--Kanade (LK) optical-flow tracking or LightGlue (LG) descriptor matching.
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
GeoFF3D is a new feed‑forward 3D reconstruction method designed for large‑scale UAV mapping. It uses a coordinate‑anchored model that predicts camera poses and dense point maps directly in a gravity‑aligned Z‑up metric frame, while a spatial large‑scale reconstruction framework (SLRF) partitions images into overlapping chunks, propagates shared‑view priors, and aggregates local reconstructions hierarchically. Across nine aerial mapping blocks, GeoFF3D achieves the best average reconstruction quality, improving F@5 from 0.829 to 0.877, and can reconstruct 2,000 images in about five minutes.
By Xiang Yang, Yongli Wang, Yunsheng Zhang
AMB3R‑SLAM is a real‑time monocular SLAM system that can reconstruct kilometer‑scale trajectories over 10,000 frames on a single consumer‑grade GPU. It combines a lightweight front‑end for low‑latency tracking with a hierarchical backend that enforces local, mid‑level, and global consistency, avoiding bundle adjustment and thus handling dynamic scenes naturally. The system also supports stereo, RGB‑D, and LiDAR inputs, achieving strong camera tracking performance and reducing absolute trajectory error by over 70% on several datasets, with sub‑meter accuracy when LiDAR is added.
By Hengyi Wang, Lourdes Agapito