RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.
By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu
arXiv:2606. 29600v1 Announce Type: cross Abstract: A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces.
By Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, Matthew Johnson-Roberson, Xiaonan Huang
arXiv:2607. 20326v1 Announce Type: cross Abstract: RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available.
By Xuchen Zhu, Yajuan Wei, Shuang Hao, Jiwei Jiang, Guanxiang Mao, Fang Ren
RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice, failures or occlusions of surveillance sensors often remove one modality.
arXiv:2608. 07579v1 Announce Type: cross Abstract: The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only.
By Abdullah Naeem, Anav Katwal, Ayon Dey, Noman Khan, Md Tamjidul Hoque
arXiv:2608.30129v1 Announce Type: new
Abstract: This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieve...
By Bingde Liu, Wu Ran, Jinglei Zhang, Huanhuan Yuan, Chao Ma
arXiv:2608.30820v1 Announce Type: new
Abstract: Occlusion boundaries (OBs) are pixel-level image boundaries corresponding to surface visibility discontinuities caused by occlusion. Through precise bo...
By Lintao Xu, Yinghao Wang, Chenchu Rong, Xuchong Qiu, Chaohui Wang
arXiv:2503. 19947v2 Announce Type: replace-cross Abstract: Generalized metric depth understanding is critical for precise vision-guided robotics, which current state-of-the-art (SOTA) vision-encoders do not support.
By Paul Koch, J\"org Kr\"uger
arXiv:2607. 17099v1 Announce Type: cross Abstract: Recent geometric foundation models (e.
By Feng Xue, Wu Chen, Mingshuai Zhao, Guofeng Zhong, Anlong Ming, Haozhe Wang, Dianqiao Lei, Zhaowen Lin, Haiyang Zhang, Nicu Sebe
arXiv:2606. 04922v1 Announce Type: cross Abstract: Current prompt-based and adapter-based tuning of vision-language models (VLMs) is attractive for medical imaging, where clinical data sensitivity favors frozen backbones and annotations are limited.
By Tran Dinh Tien, Zhiqiang Shen
DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.
By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv:2607. 02921v1 Announce Type: cross Abstract: Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants.
By Maxwell Horton, Wei Lu, Quan Tran, Yury Astashonok, Kirmani Ahmed, Babak Damavandi, Anuj Kumar, Xiao Zhang, Seungwhan Moon