SPROUT: A Scalable Diffusion Foundation Model for Multi-Crop Plant Phenotyping
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion.
PlantC2USeg is a deep transfer‑learning framework that uses cross‑scale consistency learning and an information‑restricted decoder to improve plant point cloud segmentation. It achieves state‑of‑the‑art performance on Soybean3D and ShapeNet Part, and demonstrates strong few‑shot generalization across species and sensing conditions. The method reduces the need for large annotated datasets and lowers adaptation overhead for new plant species.
arXiv:2608. 00870v1 Announce Type: cross Abstract: Panoptic crop mapping requires both delineating individual agricultural parcels and assigning a crop type to each parcel from satellite image time series.
LeafTrackNet is a deep learning framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network to track individual leaves over time. The authors introduce CanolaTrack, a large benchmark dataset of 5,704 RGB images with 31,840 annotated leaf instances from 184 canola plants. When evaluated without prior fine‑tuning, LeafTrackNet outperforms existing methods on CanolaTrack, KOMATSUNA, and MSU‑PID datasets, achieving HOTA scores of 88.03, 87.33, and 74.20 respectively.
arXiv:2608.21254v1 Announce Type: cross Abstract: Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reduci...
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.