arXiv Computer Vision

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.

Hugging Face Trending Papers
Jul 2

DL-VINS-Factory: A Modular Framework for Learned Visual Front-Ends in Visual-Inertial SLAM

Deep-learning features excel in visual matching, yet their practical value in tightly coupled visual-inertial SLAM (VI-SLAM) remains insufficiently characterized. We present DL-VINS-Factory, a unified framework that integrates learned feature extractors (ALIKED, RaCo, SuperPoint, XFeat) with either Lucas--Kanade (LK) optical-flow tracking or LightGlue (LG) descriptor matching.

arXiv Computer Vision
Sep 21

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.

By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv Computer Vision
Sep 24

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

DAVIO is a dense monocular‑inertial SLAM system that leverages a single multi‑view depth model (Depth Anything 3) for both initialization and mapping. It starts up quickly by solving a feature‑free linear system from a five‑image window and IMU pre‑integration, then uses a VIO filter whose metric poses condition the depth model during tracking. The system corrects residual scale along viewing rays, preserves metric baselines, and refines the map with a gravity‑preserving sub‑map graph, achieving earlier start‑up, lower localization error, and more accurate dense maps than state‑of‑the‑art feed‑forward mappers on both EuRoC and building‑scale ORI datasets.

By Jaafar Mahmoud, Arthur Movsesyan, Mikhail Iumanov, Sergey Kolyubin
arXiv Computer Vision
Sep 24

RoadOcc Learns When to Persist, Transport, or Refresh Memory for Roadside Occupancy Prediction

RoadOcc is a new method for roadside occupancy prediction that learns to route information among three memory sources: Persist (fixed-coordinate history), Transport (velocity-addressed history), and Refresh (current evidence). It employs dynamic-aware cross‑attention, multi‑scale voxel velocity estimation, and velocity‑guided dynamic sparse fusion to combine these sources efficiently. On the InfraOcc dataset, RoadOcc achieves 65.29 mIoU and 32.37 dynamic mIoU, outperforming the previous STCOcc baseline by significant margins.

By Xiaokai Bai, Lei Yang, Songkai Wang, Lianqing Zheng, Si-Yuan Cao, Hui-liang Shen
arXiv Computer Vision
6d ago

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.

By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
arXiv Computer Vision
Sep 18

Online Adaptation of Visual Odometry Frontends with Image-Conditioned Reinforcement Learning

The paper introduces a visual odometry frontend that automatically and continuously adapts its parameters using an image-conditioned reinforcement learning policy. The policy selects key tuning values—FAST detection threshold, KLT patch size, and RANSAC rejection threshold—based on a lightweight image embedding and frontend statistics, with a privileged critic aiding training. Trained on synthetic data, the approach transfers zero‑shot to real-world benchmarks, improving the tracking‑computation trade‑off by up to 8% in accuracy and 57% in runtime compared to static configurations.

By Simone Nascivera, Leonard Bauersfeld, Jeff Delaune, Davide Scaramuzza