Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical...
The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.
By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.
By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.
The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.
By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles
arXiv:2608. 08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding.
By Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach