AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.
By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.
arXiv:2508. 17117v3 Announce Type: replace-cross Abstract: Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis.
By Syed Nazmus Sakib, Nafiul Haque, Mohammad Zabed Hossain, Shifat E. Arman
arXiv:2608. 08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding.
By Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.
By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation.
GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.
By Ashutosh Gandhe, Anupam Rawat, Geet Sethi, Kabir Nasiruddin, Madhav Kotecha, Panav Shah, Rakshit Sawarn, Soumitra Nayak
arXiv:2606. 10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery.
By Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical...
The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.
By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
arXiv:2607. 21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions.
By Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee, Paul Liang, Siva Reddy, Verna Dankers
arXiv:2605.17949v2 Announce Type: replace
Abstract: Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the...
By Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Lang Sun, Jiaqi Liu, Jiashun Zhu, Jiasen Hu, Xu Na, Bo Yang