Hugging Face Trending Papers

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

Read the original on Hugging Face Trending Papers →

Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existing models are limited to a single scale and exhibit considerable performance degradation in complex driving environments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 7

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

The paper introduces FlexDepth, a family of self‑supervised monocular depth estimation models designed for robust driving perception. FlexDepth uses a two‑stage static‑dynamic decoupled training strategy and a Scale‑Driven Decoder that selects components based on scale size, enabling efficient feature fusion and high‑precision depth maps. Experiments on driving benchmarks show state‑of‑the‑art performance across arbitrary scales with minimal computational cost, with the smallest model (Flex‑Nano) achieving 37.6 FPS on mobile devices.

By Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan
Hugging Face Trending Papers
Jul 9

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively within single-domain, self-supervised paradigms, failing silently under domain shift.

arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv Computer Vision
6d ago

Self-Supervised Perceptually Interpretable Monocular Depth Estimation

The paper introduces PIMDE, a self‑supervised monocular depth estimation framework that decomposes input images into perceptual feature maps, each encoding a specific visual cue. Separate depth branches process these maps to produce individual depth estimates, which are then fused explicitly. Experiments on the KITTI benchmark show that PIMDE matches the accuracy of existing self‑supervised methods while offering clearer insight into how each perceptual cue contributes to depth prediction.

By Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis