arXiv Computer Vision

Robust Structureless Monocular Visual Inertial Initialization Exploiting Line Features and Vanishing Points

The paper introduces SLIM-init, a monocular visual‑inertial initialization method that avoids explicit 3D reconstruction by using 2D line features and vanishing points. It provides robust rotation constraints through line‑derived vanishing points and improves translation estimation with a line epipolar residual and a line‑normal projection residual. Experiments on public benchmarks and custom degenerate‑motion sequences show that SLIM-init achieves higher accuracy and robustness than existing initialization techniques.

arXiv Computer Vision
Sep 11

DefVINS: Visual-Inertial Odometry for Deformable Scenes

DefVINS is a visual‑inertial odometry pipeline tailored for deformable scenes, breaking the rigidity assumption of traditional VIO. It decomposes the odometry state into a rigid, IMU‑anchored component and a non‑rigid scene warp using an embedded deformation graph. The authors also introduce VIMandala, the first real‑world benchmark with ground‑truth camera poses for deformable VIO, and extend the synthetic Drunkard’s benchmark with inertial data, demonstrating that DefVINS outperforms both rigid and non‑rigid baselines.

By Samuel Cerezo, Javier Civera
arXiv Computer Vision
Sep 24

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

DAVIO is a dense monocular‑inertial SLAM system that leverages a single multi‑view depth model (Depth Anything 3) for both initialization and mapping. It starts up quickly by solving a feature‑free linear system from a five‑image window and IMU pre‑integration, then uses a VIO filter whose metric poses condition the depth model during tracking. The system corrects residual scale along viewing rays, preserves metric baselines, and refines the map with a gravity‑preserving sub‑map graph, achieving earlier start‑up, lower localization error, and more accurate dense maps than state‑of‑the‑art feed‑forward mappers on both EuRoC and building‑scale ORI datasets.

By Jaafar Mahmoud, Arthur Movsesyan, Mikhail Iumanov, Sergey Kolyubin
arXiv Computer Vision
Sep 25

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.

By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
arXiv Computer Vision
Sep 18

Monocular Visual Odometry without Calibration or Test-time Optimization

The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.

By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv Computer Vision
Sep 18

GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction

The paper introduces GRF-Recon, a framework for stable and scalable feed-forward 3D reconstruction from long monocular image sequences. It combines coarse-to-fine trajectory alignment, lightweight geometric prior injection via LoRA adaptation, and a hybrid-weight sparse ray-field optimization to refine local point clouds while enforcing cross-frame consistency. An efficient trajectory stitching strategy with joint ray-error optimization further reduces accumulated drift, achieving competitive trajectory accuracy compared to SLAM systems while maintaining globally consistent reconstructions in large-scale scenarios.

By Enpeng Li, Yunzhou Zhang, Zhiyao Zhang, Dexuan Lyu, Chenyu Wang, Chiyuan Cui, Cheng Cheng
arXiv Machine Learning
Aug 27

Minimalist Visual Inertial Odometry

The paper introduces a minimalist visual-inertial odometry system that uses only four downward-facing photodiodes with optical Gabor masks and an IMU to estimate motion for differential-drive robots. By jointly optimizing mask parameters and a Temporal Convolutional Network in a physically-grounded simulator, the model decodes speed from the photodiode signals and combines it with IMU angular speed to produce a continuous planar trajectory. Experiments on a prototype robot across indoor and outdoor terrains show that the system closely follows reference trajectories without real-world fine-tuning.

By Francesco Pasti, Jeremy Klotz, Nicola Bellotto, Shree K. Nayar
arXiv Computer Vision
5d ago

Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting

The paper introduces a reliability-regulated trajectory optimization framework for progressive COLMAP‑free 3D Gaussian Splatting (3DGS). It uses a self‑supervised bidirectional cycle‑consistency mechanism to control camera trajectory estimation through forward motion propagation and retrospective trajectory correction, thereby reducing error compounding without external priors. Experiments on Tanks and Temples and CO3D‑V2 demonstrate improved camera trajectory accuracy and novel‑view rendering quality compared to existing unposed baselines.

By Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu, Hanjie Ma, Zhen Ye, Mingfeng Jiang