arXiv Computer Vision

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

Con-DSO introduces a consistency-aware RGB‑D direct sparse odometry framework that learns pixel‑level photometric and geometric uncertainty from adjacent RGB‑D frame pairs. The network predicts uncertainties that are converted into pairwise quality scores, guiding support‑pixel selection and forming a host‑side quality prior for keyframe tracking. Experiments on five public benchmarks show that this approach reduces absolute trajectory error by over 20% on ICL‑NUIM and by 50–80% on other datasets, improving robustness in challenging environments.

arXiv Computer Vision
Sep 25

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.

By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
arXiv Computer Vision
Sep 18

Monocular Visual Odometry without Calibration or Test-time Optimization

The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.

By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv Computer Vision
1d ago

PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

The paper introduces PointVGGT, a zero‑shot framework for multiview RGB‑D point cloud registration that replaces the traditional pairwise‑then‑global pipeline. It employs a foundation‑then‑refinement paradigm, using visual geometry foundation models to directly recover metrically consistent global poses and then refining them with voxelized spatial hashing and IRLS‑based bundle adjustment. Experiments on indoor, object‑centric, and outdoor datasets demonstrate superior registration accuracy and computational efficiency without any training.

By Haobo Jiang, Liang Yu, Jianmin Zheng
arXiv Computer Vision
Sep 21

SFVO: Decoupled Confidence-Guided Stereo-Flow Visual Odometry with Bidirectional PnP

SFVO is a stereo visual‑odometry framework that leverages pretrained stereo‑matching and optical‑flow models to obtain dense stereo and temporal correspondences. Rather than learning pose directly from images, it maps these correspondences into geometric constraints and predicts trustworthy points using decoupled confidence maps for rotation and translation. Experiments on both outdoor and indoor datasets show that SFVO delivers robust, accurate pose estimation with strong generalization, and the authors plan to release the code.

By Kai Zhang, Guoyang Zhao, Jun Ma
arXiv Computer Vision
Sep 28

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.

By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow