arXiv:2606. 13509v1 Announce Type: cross Abstract: Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline.
By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
Hydra introduces a marker‑free RGB‑D hand‑eye calibration method that leverages a novel ICP algorithm with a robust point‑to‑plane objective on a Lie algebra. Experiments on three serial manipulators and two RGB‑D cameras show that with only three random robot configurations the method achieves about 90% successful calibrations, 2–3× faster convergence to the global optimum, and 2 orders of magnitude faster convergence time (0.8 ± 0.4 s) compared to other marker‑free baselines. The approach delivers improved accuracy (5 mm in task space versus 7 mm for classical methods) while remaining marker‑free, and the authors provide an open‑source dataset, code, and ROS 2 integration.
By Martin Huber, Huanyu Tian, Christopher E. Mower, Lucas-Raphael M\"uller, S\'ebastien Ourselin, Christos Bergeles, Tom Vercauteren
arXiv:2609.10082v1 Announce Type: cross
Abstract: Accurate camera intrinsic calibration is fundamental to robot perception, and the accuracy depends on the quality of the collected images. However, e...
By Xiangcheng Hu
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Language-Action (VLA) policies tolerate such camera variations only when the camera extrinsics are explicitly provided, making them fragile and hard to use especially when view robustness is critical.
arXiv:2607. 05396v1 Announce Type: cross Abstract: Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios.
By Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Shijian Lu, Gongjie Zhang, Ran Xu
arXiv:2610.01744v1 Announce Type: new
Abstract: Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented ro...
By Wonguen Cho, Junhoo Lee, Nojun Kwak
The paper introduces a minimalist visual-inertial odometry system that uses only four downward-facing photodiodes with optical Gabor masks and an IMU to estimate motion for differential-drive robots. By jointly optimizing mask parameters and a Temporal Convolutional Network in a physically-grounded simulator, the model decodes speed from the photodiode signals and combines it with IMU angular speed to produce a continuous planar trajectory. Experiments on a prototype robot across indoor and outdoor terrains show that the system closely follows reference trajectories without real-world fine-tuning.
By Francesco Pasti, Jeremy Klotz, Nicola Bellotto, Shree K. Nayar
Robust dynamic object detection and tracking are essential for enabling robots to operate safely and effectively alongside humans in complex environments such as construction sites. While LiDAR-based SLAM and occupancy grid methods offer viable solutions for detecting and tracking motion, many state-of-the-art 3D vision approaches rely heavily on pre-trained neural networks and require additional post-processing to identify moving objects.
The paper introduces TempLoc, a Temporal‑aware Localization framework that improves outdoor LiDAR relocalization by leveraging spatio‑temporal consistency across scans. It first predicts point‑wise global coordinates with uncertainties, then estimates inter‑frame correspondences using an attention‑based Prior Coordinate Generation module, and finally fuses these predictions in an uncertainty‑guided manner to produce a more accurate global 6‑DoF pose. Experiments on the NCLT and Oxford RobotCar datasets show that TempLoc significantly outperforms existing state‑of‑the‑art methods.
By Minghang Zhu, Zhijing Wang, Yuxin Guo, Chen Liu, Yongshu Huang, Wen Li, Sheng Ao, Cheng Wang
The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.
By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv:2609.25746v1 Announce Type: cross
Abstract: ICP-based 3D Gaussian Splatting (3DGS) SLAM tracks in real time by registering incoming frames against map Gaussians, using each primitive's covarian...
By Edward Beng Wai Tan, Siew-Kei Lam
The paper introduces an on-the-fly homography calibration system for multi-camera tracking that starts from coarse manual homographies and refines them using a centroid-based projection optimization (PO) on live detection metadata. PO continuously aligns ground-plane geometry without adding computational latency, enabling the system to adapt automatically to camera movements or environmental changes. The refined geometry feeds a bird's-eye-view tracker that fuses detections and unifies trajectories across zones while maintaining privacy safety and zero overhead.
By David Voihanski, Mor Sinai, Ben Zion Bobrovsky