arXiv:2609.13777v1 Announce Type: cross
Abstract: Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, depth, bias, uncertainty, or confidence...
By Jinchang Zhang, Guoyu Lu
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv:2608.20891v1 Announce Type: cross
Abstract: We present a vision-only state estimation system for X-configuration quadcopters equipped with a canonical stereo camera pair and no inertial sensors...
By Daniel Gr{\o}nhaug, Sofie Markeset, Mathias Kolberg
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
DAVIO is a dense monocular‑inertial SLAM system that leverages a single multi‑view depth model (Depth Anything 3) for both initialization and mapping. It starts up quickly by solving a feature‑free linear system from a five‑image window and IMU pre‑integration, then uses a VIO filter whose metric poses condition the depth model during tracking. The system corrects residual scale along viewing rays, preserves metric baselines, and refines the map with a gravity‑preserving sub‑map graph, achieving earlier start‑up, lower localization error, and more accurate dense maps than state‑of‑the‑art feed‑forward mappers on both EuRoC and building‑scale ORI datasets.
By Jaafar Mahmoud, Arthur Movsesyan, Mikhail Iumanov, Sergey Kolyubin