arXiv Computer Vision By Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

Read the original on arXiv Computer Vision →

AgroBench is a reproducible benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel‑level crop time series. The data generation pipeline fuses USDA yield data with land cover masks, Sentinel‑2 and Sentinel‑1 imagery, climatic variables, and terrain information to produce multimodal sequences for individual crop pixels across the growing season. The benchmark includes over 13 million observations from 788,654 crop pixels, covering 5,107 county‑year combinations for five major U.S. crops from 2017 to 2024, and establishes a Leave‑One‑Year‑Out evaluation protocol with baseline machine learning results.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Sep 17

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.

arXiv Computer Vision
Sep 18

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.

By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
Hugging Face Trending Papers
Jul 22

Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data

Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics and per-plant measurement labels are scarce. In this paper, we introduce a novel, annotated image time-series dataset of 691 sweet pepper plants monitored over two growing seasons, comprising 4837 images with per-plant fruit counts categorized by maturity.