arXiv:2610.03283v1 Announce Type: cross
Abstract: The ability of a device to localize itself within its surroundings is a fundamental prerequisite for spatial computing. Visual-inertial odometry (VIO...
By Patrick Wolf, Mateo de Mayo, Daniel Cremers
arXiv:2609.23974v1 Announce Type: new
Abstract: Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through vis...
By Boxun Hu, Jiawei Ge, Axel Krieger, Peng Wang, Tinoosh Mohsenin
The paper examines how multi‑camera streaming perception systems perform on heterogeneous edge platforms that share resources with other workloads. Using two end‑to‑end pipelines on a single GPU–NPU platform, the authors show that isolated single‑stream evaluations can mislead deployment decisions: while the GPU pipeline appears superior in isolation, GPU‑local contention causes deadline misses that make detections stale and can reverse the preferred placement. The study finds that the NPU pipeline, though less accurate for small and medium objects, nearly matches the GPU on large objects, and that under high contention the best placement shifts from All‑GPU to All‑NPU, achieving a 5.2× improvement in worst‑stream sAP. The authors argue that evaluation metrics should include contention sweeps, deadline‑miss rates, and worst‑stream sAP in addition to mean sAP to capture severe single‑stream degradation.
By Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee
arXiv:2608. 14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure.
By Haibo HU, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
arXiv:2607.08098v2 Announce Type: replace
Abstract: Event cameras are increasingly adopted in embodied perception for their microsecond temporal resolution, high dynamic range, and resilience to moti...
By Linli Shi, Ruijun Zhang, Ziyun Wang
Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human...
arXiv:2608. 01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries.
By Dzmitry Malyshau
AMB3R‑SLAM is a real‑time monocular SLAM system that can reconstruct kilometer‑scale trajectories over 10,000 frames on a single consumer‑grade GPU. It combines a lightweight front‑end for low‑latency tracking with a hierarchical backend that enforces local, mid‑level, and global consistency, avoiding bundle adjustment and thus handling dynamic scenes naturally. The system also supports stereo, RGB‑D, and LiDAR inputs, achieving strong camera tracking performance and reducing absolute trajectory error by over 70% on several datasets, with sub‑meter accuracy when LiDAR is added.
By Hengyi Wang, Lourdes Agapito
arXiv:2607. 24148v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment.
By Zhuoran Song, Haozhe Jiang, Chunyu Qi, Minnan Pei, Gang Li, Xiaoyao Liang, Haibing Guan
Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.
By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
The paper introduces a minimalist visual-inertial odometry system that uses only four downward-facing photodiodes with optical Gabor masks and an IMU to estimate motion for differential-drive robots. By jointly optimizing mask parameters and a Temporal Convolutional Network in a physically-grounded simulator, the model decodes speed from the photodiode signals and combines it with IMU angular speed to produce a continuous planar trajectory. Experiments on a prototype robot across indoor and outdoor terrains show that the system closely follows reference trajectories without real-world fine-tuning.
By Francesco Pasti, Jeremy Klotz, Nicola Bellotto, Shree K. Nayar
arXiv:2608.24544v1 Announce Type: new
Abstract: Many feature-based visual-inertial odometry (VIO) systems rely on sparse feature tracking, whose accuracy and robustness directly affect state estimati...
By Renbiao Jin, Danping Zou, Wenxian Yu