The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.
By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv:2603. 28251v3 Announce Type: replace-cross Abstract: Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety.
By Weimin Liu, Qingkun Li, Jiyuan Qiu, Wenjun Wang, Joshua H. Meng
Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.
arXiv:2606. 06899v1 Announce Type: cross Abstract: Variations in illumination remain a major challenge for visual representation learning, as they induce substantial appearance changes both across and within environments.
By Lizhen Zhu, Charantej Reddy Pochimireddy, James Z Wang, Brad Wyble
RGBD20K is a new large-scale RGB‑D semantic segmentation dataset featuring 20,000 image pairs and 160 fine‑grained categories, surpassing existing benchmarks like NYUv2 and SUN RGB‑D in both scale and semantic diversity. The dataset provides high‑fidelity annotations obtained through rigorous re‑evaluation and correction of prior labels, ensuring a clean ground‑truth foundation. Additionally, the authors introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art performance across evaluated benchmarks, demonstrating the value of high‑quality multimodal information.
By Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan
RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.