The paper introduces FlexDepth, a family of self‑supervised monocular depth estimation models designed for robust driving perception. FlexDepth uses a two‑stage static‑dynamic decoupled training strategy and a Scale‑Driven Decoder that selects components based on scale size, enabling efficient feature fusion and high‑precision depth maps. Experiments on driving benchmarks show state‑of‑the‑art performance across arbitrary scales with minimal computational cost, with the smallest model (Flex‑Nano) achieving 37.6 FPS on mobile devices.
By Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan
Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively within single-domain, self-supervised paradigms, failing silently under domain shift.
arXiv:2608.13147v2 Announce Type: replace
Abstract: Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera str...
By Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li, Ya Yang, Naiyan Wang, Si Liu
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv:2609.22896v1 Announce Type: new
Abstract: Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo....
By Serin Varghese, Fabian H\"uger, Kira Maag
The paper introduces PIMDE, a self‑supervised monocular depth estimation framework that decomposes input images into perceptual feature maps, each encoding a specific visual cue. Separate depth branches process these maps to produce individual depth estimates, which are then fused explicitly. Experiments on the KITTI benchmark show that PIMDE matches the accuracy of existing self‑supervised methods while offering clearer insight into how each perceptual cue contributes to depth prediction.
By Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis