arXiv Computer Vision

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D semantic segmentation dataset featuring 20,000 image pairs and 160 fine‑grained categories, surpassing existing benchmarks like NYUv2 and SUN RGB‑D in both scale and semantic diversity. The dataset provides high‑fidelity annotations obtained through rigorous re‑evaluation and correction of prior labels, ensuring a clean ground‑truth foundation. Additionally, the authors introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art performance across evaluated benchmarks, demonstrating the value of high‑quality multimodal information.

Hugging Face Trending Papers
Sep 24

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.

arXiv AI
Sep 2

SARTM: Segment Any RGB Thermal Model with Language aided Distillation

The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.

By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv AI
Aug 3

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.

By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv AI
Aug 19

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

The paper introduces OVRSISBench, a unified benchmark for open‑vocabulary remote sensing image segmentation, and evaluates existing OVS/OVRSIS models, uncovering their shortcomings in remote sensing contexts. Leveraging insights from this evaluation, the authors propose RSKT‑Seg, a new framework featuring a Multi‑Directional Cost Map Aggregation module, an Efficient Cost Map Fusion transformer, and a Remote Sensing Knowledge Transfer module. Experiments on the benchmark demonstrate that RSKT‑Seg outperforms strong baselines by +3.8 mIoU and +5.9 mACC while achieving twice the inference speed.

By Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li