MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.
By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
arXiv:2609.36929v1 Announce Type: new
Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...
By Thai Duy Nguyen, Addison Lin Wang
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
The paper introduces PointVGGT, a zero‑shot framework for multiview RGB‑D point cloud registration that replaces the traditional pairwise‑then‑global pipeline. It employs a foundation‑then‑refinement paradigm, using visual geometry foundation models to directly recover metrically consistent global poses and then refining them with voxelized spatial hashing and IRLS‑based bundle adjustment. Experiments on indoor, object‑centric, and outdoor datasets demonstrate superior registration accuracy and computational efficiency without any training.
By Haobo Jiang, Liang Yu, Jianmin Zheng
SFVO is a stereo visual‑odometry framework that leverages pretrained stereo‑matching and optical‑flow models to obtain dense stereo and temporal correspondences. Rather than learning pose directly from images, it maps these correspondences into geometric constraints and predicts trustworthy points using decoupled confidence maps for rotation and translation. Experiments on both outdoor and indoor datasets show that SFVO delivers robust, accurate pose estimation with strong generalization, and the authors plan to release the code.
By Kai Zhang, Guoyang Zhao, Jun Ma
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza