To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation.
arXiv:2608. 08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding.
By Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
arXiv:2508. 17117v3 Announce Type: replace-cross Abstract: Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis.
By Syed Nazmus Sakib, Nafiul Haque, Mohammad Zabed Hossain, Shifat E. Arman
arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.
By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
arXiv:2609.09417v1 Announce Type: new
Abstract: Vision-language models (VLMs) show promise for agricultural classification, but zero-shot performance on disease, pest, damage, quality, and species id...
By Earl Ranario, Jared Smith, Lars Lundqvist, Urmil Jatin Chandarana
arXiv:2609.10469v1 Announce Type: new
Abstract: Automated plant disease diagnosis is increasingly deployed on farmer-held devices in regions where agronomic expertise is scarce and network connectivi...
By Md. Abdullah Mandal, Saad Ahmed, Md. Khalid Syfullah
arXiv:2608.21254v1 Announce Type: cross
Abstract: Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reduci...
By Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez, Shekhar Borah, Athresh Karanam, Erik Blasch, Prabha Sundaravadivel, Sriraam Natarajan
CoAtNet-DeepMoE is a lightweight Convolution‑Attention hybrid architecture that incorporates a DeepSeek Mixture‑of‑Experts to reduce parameters while maintaining high accuracy for tomato disease classification. The model achieves state‑of‑the‑art performance on Kaggle and PlantVillage datasets, reporting 99.80% accuracy on Kaggle and 99.83% accuracy on PlantVillage, all with only 2.47 million parameters. The source code will be released on GitHub.
By Md Nadim Mahamood, Md Arif Shahriar, Md Shafi Ud Doula, Kamrul Hasan
The paper studies how prompt quality affects vision‑language models in the agrifood domain. It evaluates Zero‑Shot Prompt Ensembling (ZPE) on CLIP and SigLIP across four datasets, showing that ZPE offers limited gains on in‑distribution data but significantly improves accuracy and calibration when the domain shifts. The authors also introduce Prompt‑based Inconsistency Detection (PID), which uses prompt disagreement to detect failures under severe domain shift, outperforming standard confidence measures.
By Andrea Morales-Garz\'on, Salvador L\'opez-Joya, Miguel L\'opez-P\'erez, Maria J. Martin-Bautista
arXiv:2603. 14342v2 Announce Type: replace-cross Abstract: Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery.
By Jiarui Zhang, Junqi Hu, Zurong Mai, Yang Liu, Yuhang Chen, Shuohong Lou, Henglian Huang, Hong Cheng, Lingyuan Zhao, Jianxi Huang, Yutong Lu, Haohuan Fu, Juepeng Zheng
The paper introduces the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), which fuses decision-level outputs from EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by multimodal large language models Gemma 4 E4B and Qwen3.5 4B to produce explainable plant disease diagnoses. Evaluated on 14,364 images from PlantDoc and two Cornell robotic field datasets, the framework achieves up to 99.3% accuracy, with Gemma improving PlantDoc accuracy from 63.9% to 68.5% and demonstrating low critical‑risk error. The results highlight the potential of MLLM arbitration for reliable, explainable agricultural AI under real‑world field conditions.
By Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical...