Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics and per-plant measurement labels are scarce. In this paper, we introduce a novel, annotated image time-series dataset of 691 sweet pepper plants monitored over two growing seasons, comprising 4837 images with per-plant fruit counts categorized by maturity.
LeafTrackNet is a deep learning framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network to track individual leaves over time. The authors introduce CanolaTrack, a large benchmark dataset of 5,704 RGB images with 31,840 annotated leaf instances from 184 canola plants. When evaluated without prior fine‑tuning, LeafTrackNet outperforms existing methods on CanolaTrack, KOMATSUNA, and MSU‑PID datasets, achieving HOTA scores of 88.03, 87.33, and 74.20 respectively.
By Shanghua Liu, Majharulislam Babor, Christoph Verduyn, Breght Vandenberghe, Bruno Betoni Parodi, Cornelia Weltzien, Marina M. -C. H\"ohne
AgroBench is a reproducible benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel‑level crop time series. The data generation pipeline fuses USDA yield data with land cover masks, Sentinel‑2 and Sentinel‑1 imagery, climatic variables, and terrain information to produce multimodal sequences for individual crop pixels across the growing season. The benchmark includes over 13 million observations from 788,654 crop pixels, covering 5,107 county‑year combinations for five major U.S. crops from 2017 to 2024, and establishes a Leave‑One‑Year‑Out evaluation protocol with baseline machine learning results.
By Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat
The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.
By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles
arXiv:2603.27519v4 Announce Type: replace
Abstract: Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, orga...
By Shuai Xiang, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu, Yinqiang Zheng, Wei Guo
arXiv:2608.30088v1 Announce Type: new
Abstract: Accurate detection of tomato growth stages is essential for stage-specific greenhouse management and precision agriculture. In Bhutan, greenhouse culti...
By Sherab Gocha, Sou Nobukawa
PlantC2USeg is a deep transfer‑learning framework that uses cross‑scale consistency learning and an information‑restricted decoder to improve plant point cloud segmentation. It achieves state‑of‑the‑art performance on Soybean3D and ShapeNet Part, and demonstrates strong few‑shot generalization across species and sensing conditions. The method reduces the need for large annotated datasets and lowers adaptation overhead for new plant species.
By Yu Tian, Xintong Jiang, Jan Franklin Adamowski, Shiv O. Prasher, Shangpeng Sun
arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.
By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.
arXiv:2609.09062v1 Announce Type: new
Abstract: We present a real-world case study of multi-task learning (MTL) for temporal process modeling from limited data with temporally sparse labels. Specific...
By Aseem Saxena, Paola Pes\'antez-Cabrera, Jonathan Magby, Markus Keller, Alan Fern
arXiv:2609.10469v1 Announce Type: new
Abstract: Automated plant disease diagnosis is increasingly deployed on farmer-held devices in regions where agronomic expertise is scarce and network connectivi...
By Md. Abdullah Mandal, Saad Ahmed, Md. Khalid Syfullah
arXiv:2609.21059v1 Announce Type: cross
Abstract: Plant growth and agricultural production form the foundation of a country's sustainable development and directly impact human livelihoods. Recent adv...
By Longchao Da, Xiaoou Liu, Xingjian Li, Lirong Xiang, Hua Wei