arXiv:2606. 17020v1 Announce Type: cross Abstract: Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored.
By Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang, Yuxian Dong, Chengyin Hu, Fengyu Zhang, Yiwei Wei, Jiujiang Guo
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared images provide distinctive cues, including thermal intensity structures, object boundaries, and illumination-invariant scene features, which can enrich visual-language learning beyond conventional RGB observations.
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
The paper explores using generative models to translate RGB UAV images into synthetic infrared (IR) images for training vehicle detectors in domains where real IR data is scarce. Various translators—supervised GANs, ControlNet-based diffusion models, and LoRA-ed foundation models—were trained on paired RGB-IR datasets and applied to unseen target datasets to generate synthetic IR data. The synthetic IR images, especially those produced by Stable Diffusion 3.5 with ControlNet, significantly improved detection performance on unseen IR test sets, outperforming RGB and grayscale baselines and narrowing the gap to real IR data.
By Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden, Elfi I. S. Hofmeijer, Sebastiaan P. Snel, Klamer Schutte, Friso G. Heslinga
arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.
By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.
By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed