AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.
By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion.
arXiv:2607. 03245v1 Announce Type: new Abstract: High-throughput plant phenotyping generates valuable data that often remains trapped in unstructured text and isolated RGB images.
By Jayant Ghadge, Soumyashree Kar, Surya S. Durbha
The paper introduces AGS-PlantSeg, a few‑shot 3D plant organ segmentation method that uses the frozen Utonia foundation model and Adaptive Granularity Selection (AGS) to dynamically choose optimal spatial granularity for each plant. By extracting tailored geometric features for a lightweight MLP head, AGS-PlantSeg achieves superior cross‑species generalization, reaching an average mIoU of 88.9% and outperforming fixed‑granularity baselines by 2.5 points across PLANesT‑3D, Pheno4D, and Crops3D datasets. The approach requires minimal annotated data yet competes with fully supervised, plant‑specific architectures.
The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.
By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee