arXiv Machine Learning

GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry

GINIO is a geometric SO(3)-equivariant interface designed for neural inertial odometry that ensures learned measurements transform consistently under any IMU mounting convention. It predicts motion measurements and uncertainties that obey vector and tensor transformation laws, and introduces Last-Frame Alignment to enable efficient sensor-frame learning equivalent to world-frame training. The interface is instantiated in several architectures—filter-connected NIO, AirIO-style recurrent aerial prediction, EqNIO-style full-SO(3) canonicalization, and ResNet-style temporal backbones—achieving significant accuracy and efficiency gains across multiple benchmarks.

arXiv Computer Vision
Sep 21

Refining Ground Truth Poses in Autonomous Driving Datasets via Neural Rendering

arXiv:2504.15776v2 Announce Type: replace Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...

By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv Computer Vision
3d ago

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.

By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
Hugging Face Trending Papers
Jun 23

REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.

arXiv Computer Vision
Sep 24

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

DAVIO is a dense monocular‑inertial SLAM system that leverages a single multi‑view depth model (Depth Anything 3) for both initialization and mapping. It starts up quickly by solving a feature‑free linear system from a five‑image window and IMU pre‑integration, then uses a VIO filter whose metric poses condition the depth model during tracking. The system corrects residual scale along viewing rays, preserves metric baselines, and refines the map with a gravity‑preserving sub‑map graph, achieving earlier start‑up, lower localization error, and more accurate dense maps than state‑of‑the‑art feed‑forward mappers on both EuRoC and building‑scale ORI datasets.

By Jaafar Mahmoud, Arthur Movsesyan, Mikhail Iumanov, Sergey Kolyubin
arXiv Computer Vision
6d ago

Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces

The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.

By Edward Beng Wai Tan, Siew-Kei Lam
arXiv Computer Vision
Sep 18

Monocular Visual Odometry without Calibration or Test-time Optimization

The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.

By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv Computer Vision
Sep 11

DefVINS: Visual-Inertial Odometry for Deformable Scenes

DefVINS is a visual‑inertial odometry pipeline tailored for deformable scenes, breaking the rigidity assumption of traditional VIO. It decomposes the odometry state into a rigid, IMU‑anchored component and a non‑rigid scene warp using an embedded deformation graph. The authors also introduce VIMandala, the first real‑world benchmark with ground‑truth camera poses for deformable VIO, and extend the synthetic Drunkard’s benchmark with inertial data, demonstrating that DefVINS outperforms both rigid and non‑rigid baselines.

By Samuel Cerezo, Javier Civera
arXiv Computer Vision
Sep 23

Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation

The paper presents a new monocular spacecraft pose estimation model that achieves the lowest reported mean rotation errors on the SPEED+ lightbox and sunlamp test sets. By replacing smaller encoders with a large self‑supervised ViT foundation model (DINOv3) and scaling up to 840 M parameters, the authors improve accuracy from 300 M to 840 M parameters without saturation. The 840 M model also runs on a Jetson Orin NX 16 GB with 133.8 ms per crop and 32.0 W power draw, demonstrating embedded inference feasibility while training solely on synthetic data.

By John Church, Vazghen Nikolian