arXiv Machine Learning

Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

The paper presents a configurable multi‑stage vision pipeline for Farmer.Chat, a farm advisory service that processes farmer‑submitted crop photos. It splits the task into a quality gate, a crop detector, and a disease/pest detector, offering two routes: a single fine‑tuned vision‑language model and a set of small specialist models. The new pipeline improves crop accuracy to 95.41% and provides a fast MobileNetV3 quality gate, while retaining the ability to request better photos and answer all queries in one call.

arXiv AI
Sep 12

Can Edge-Deployable Vision-Language Models Identify Species?

arXiv:2609. 11916v1 Announce Type: new Abstract: Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification.

By William Zhou, Mayukha Siripuram, Xiao Yan, Ziqi Liu, Yi Ding
arXiv Computer Vision
Aug 27

Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

The paper introduces the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), which fuses decision-level outputs from EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by multimodal large language models Gemma 4 E4B and Qwen3.5 4B to produce explainable plant disease diagnoses. Evaluated on 14,364 images from PlantDoc and two Cornell robotic field datasets, the framework achieves up to 99.3% accuracy, with Gemma improving PlantDoc accuracy from 63.9% to 68.5% and demonstrating low critical‑risk error. The results highlight the potential of MLLM arbitration for reliable, explainable agricultural AI under real‑world field conditions.

By Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee
arXiv Machine Learning
Aug 27

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

CropCop is a closed‑set plant‑health recognition system covering 120 operational classes, built from a rigorously audited dataset of 109,107 images after removing 3,233 duplicate relationships. The model, based on a fine‑tuned DINOv3 ConvNeXt‑Tiny, achieves 98.51% accuracy and 96.87% macro‑F1 on a locked internal test, while a quantised MobileNetV4 variant reaches 98.46% accuracy and 96.23% macro‑F1 in a 22.60 MiB runtime artifact. Validation‑only post‑training quantisation and a compact ExecuTorch/XNNPACK PTE ensure high fidelity between the trained model and its deployed form, with minimal decision changes between the INT8 graph and the final artifact.

By Rana Muhammad Ahmed, Sabahat Abbas
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li