arXiv:2609.13777v1 Announce Type: cross
Abstract: Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, depth, bias, uncertainty, or confidence...
By Jinchang Zhang, Guoyu Lu
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv:2608.20891v1 Announce Type: cross
Abstract: We present a vision-only state estimation system for X-configuration quadcopters equipped with a canonical stereo camera pair and no inertial sensors...
By Daniel Gr{\o}nhaug, Sofie Markeset, Mathias Kolberg
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.
DAVIO is a dense monocular‑inertial SLAM system that leverages a single multi‑view depth model (Depth Anything 3) for both initialization and mapping. It starts up quickly by solving a feature‑free linear system from a five‑image window and IMU pre‑integration, then uses a VIO filter whose metric poses condition the depth model during tracking. The system corrects residual scale along viewing rays, preserves metric baselines, and refines the map with a gravity‑preserving sub‑map graph, achieving earlier start‑up, lower localization error, and more accurate dense maps than state‑of‑the‑art feed‑forward mappers on both EuRoC and building‑scale ORI datasets.
By Jaafar Mahmoud, Arthur Movsesyan, Mikhail Iumanov, Sergey Kolyubin
arXiv:2601. 03040v2 Announce Type: replace-cross Abstract: A fundamental requirement for full autonomy is the ability to sustain accurate navigation in the absence of external data, such as GNSS signals or visual information.
By Arup Kumar Sahoo, Itzik Klein
The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.
By Edward Beng Wai Tan, Siew-Kei Lam
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
DefVINS is a visual‑inertial odometry pipeline tailored for deformable scenes, breaking the rigidity assumption of traditional VIO. It decomposes the odometry state into a rigid, IMU‑anchored component and a non‑rigid scene warp using an embedded deformation graph. The authors also introduce VIMandala, the first real‑world benchmark with ground‑truth camera poses for deformable VIO, and extend the synthetic Drunkard’s benchmark with inertial data, demonstrating that DefVINS outperforms both rigid and non‑rigid baselines.
By Samuel Cerezo, Javier Civera
arXiv:2607. 02561v1 Announce Type: cross Abstract: Consumer depth sensors such as the LiDAR scanner on recent iPhones provide metric range, but their useful range is short and their returns are sparse.
By Jinwen Wen
The paper presents a new monocular spacecraft pose estimation model that achieves the lowest reported mean rotation errors on the SPEED+ lightbox and sunlamp test sets. By replacing smaller encoders with a large self‑supervised ViT foundation model (DINOv3) and scaling up to 840 M parameters, the authors improve accuracy from 300 M to 840 M parameters without saturation. The 840 M model also runs on a Jetson Orin NX 16 GB with 133.8 ms per crop and 32.0 W power draw, demonstrating embedded inference feasibility while training solely on synthetic data.
By John Church, Vazghen Nikolian