arXiv:2609.09417v1 Announce Type: new
Abstract: Vision-language models (VLMs) show promise for agricultural classification, but zero-shot performance on disease, pest, damage, quality, and species id...
By Earl Ranario, Jared Smith, Lars Lundqvist, Urmil Jatin Chandarana
arXiv:2609.10469v1 Announce Type: new
Abstract: Automated plant disease diagnosis is increasingly deployed on farmer-held devices in regions where agronomic expertise is scarce and network connectivi...
By Md. Abdullah Mandal, Saad Ahmed, Md. Khalid Syfullah
arXiv:2609. 11916v1 Announce Type: new Abstract: Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification.
By William Zhou, Mayukha Siripuram, Xiao Yan, Ziqi Liu, Yi Ding
arXiv:2609.25040v1 Announce Type: cross
Abstract: Banana crop diseases threaten food security across the world, yet field diagnosis remains difficult because of limited expert access and visual simil...
By Sangam Kumar Jena, Pandarasamy Arjunan
The paper introduces the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), which fuses decision-level outputs from EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by multimodal large language models Gemma 4 E4B and Qwen3.5 4B to produce explainable plant disease diagnoses. Evaluated on 14,364 images from PlantDoc and two Cornell robotic field datasets, the framework achieves up to 99.3% accuracy, with Gemma improving PlantDoc accuracy from 63.9% to 68.5% and demonstrating low critical‑risk error. The results highlight the potential of MLLM arbitration for reliable, explainable agricultural AI under real‑world field conditions.
By Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee
arXiv:2608.21762v1 Announce Type: cross
Abstract: Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image...
By Jinchang Zhu, Rong Fu, Yi Ding, Chenghao Wu, Ying Liu, Menglin Yang
arXiv:2607. 10666v1 Announce Type: cross Abstract: Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available.
By Shubham Rao
arXiv:2608.20663v1 Announce Type: new
Abstract: Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation...
By Pushuo Wang (Shenyang Institute of Technology)
arXiv:2608. 08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding.
By Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
arXiv:2607. 22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting.
By M M Asif Ferdous
CropCop is a closed‑set plant‑health recognition system covering 120 operational classes, built from a rigorously audited dataset of 109,107 images after removing 3,233 duplicate relationships. The model, based on a fine‑tuned DINOv3 ConvNeXt‑Tiny, achieves 98.51% accuracy and 96.87% macro‑F1 on a locked internal test, while a quantised MobileNetV4 variant reaches 98.46% accuracy and 96.23% macro‑F1 in a 22.60 MiB runtime artifact. Validation‑only post‑training quantisation and a compact ExecuTorch/XNNPACK PTE ensure high fidelity between the trained model and its deployed form, with minimal decision changes between the INT8 graph and the final artifact.
By Rana Muhammad Ahmed, Sabahat Abbas
arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.
By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li