arXiv Computer Vision

RealOOB: A Definition-Consistent Real-World Oriented Occlusion Boundary Benchmark

arXiv AI
Aug 19

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

PXDepth is a monocular depth estimation model that separates global context modeling from pixel-level depth prediction. It uses a large-patch Vision Transformer to capture scene context and a pixel-space predictor with Context‑Modulated Pixel Transformer blocks to preserve high‑resolution spatial details. The approach maintains fine structures and sharp boundaries while achieving competitive global depth accuracy in zero‑shot benchmarks.

By Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
arXiv Machine Learning
4d ago

Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

The paper introduces a depth‑aware pothole detection framework that fuses RGB‑D sensor data and evaluates five architectures—YOLOv8n, YOLOv8nSeg, YOLOv9t, RTDETRL, and RTDETRX—on the PothRGBD dataset. YOLOv8nSeg achieves the highest detection performance (mAP@50 = 0.9556, mAP@50_95 = 0.6758) and the most accurate depth estimate (2.96 cm), while YOLOv8n offers the fastest inference (3.6 ms) and RTDETRX delivers the highest detection confidence (92.70 %). The study also shows that even after RANSAC orthorectification, bounding‑box models overestimate pothole depth by 0.16–0.21 cm, indicating a structural bias rather than a calibration error.

By Md Monjurul Ahsan Prodhan, Md Nour Hossain