The paper introduces FlexDepth, a family of self‑supervised monocular depth estimation models designed for robust driving perception. FlexDepth uses a two‑stage static‑dynamic decoupled training strategy and a Scale‑Driven Decoder that selects components based on scale size, enabling efficient feature fusion and high‑precision depth maps. Experiments on driving benchmarks show state‑of‑the‑art performance across arbitrary scales with minimal computational cost, with the smallest model (Flex‑Nano) achieving 37.6 FPS on mobile devices.
By Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan
Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existing models are limited to a single scale and exhibit considerable performance degradation in complex driving environments.
The paper presents a lightweight monocular vision pipeline for extracting sidewalk paths on low‑power embedded micromobility platforms. It evolves through three design iterations—from a skeleton‑graph baseline to a distance‑transform corridor planner and finally to a compact image‑space architecture—using a SegFormer‑B0 student model trained with semi‑supervised pseudo‑labels. The final system achieves high segmentation accuracy (IoU 0.946) and fast planning (under 50 ms per frame) while reducing temporal instability and increasing template‑path availability across real campus sequences.
By Lkhanaajav Mijiddorj, Yang Yan, Tyler Beringer, Bilguunzaya Mijiddorj, Alex N. Ho, Bin Xu, Binbin Weng
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.
By Edward Beng Wai Tan, Siew-Kei Lam
arXiv:2607. 07885v1 Announce Type: cross Abstract: Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical.
By Erik Jagnandan, Mulugeta Haile, Gregory Barber, Pratik Chaudhari