Hugging Face Trending Papers

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

Read the original on Hugging Face Trending Papers →

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 18

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.

By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
arXiv AI
Jun 9

AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning

arXiv:2603. 14342v2 Announce Type: replace-cross Abstract: Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery.

By Jiarui Zhang, Junqi Hu, Zurong Mai, Yang Liu, Yuhang Chen, Shuohong Lou, Henglian Huang, Hong Cheng, Lingyuan Zhao, Jianxi Huang, Yutong Lu, Haohuan Fu, Juepeng Zheng
Hugging Face Trending Papers
Jul 7

Vision as Unified Multimodal Generation

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks.

arXiv AI
Aug 3

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.

By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv Computer Vision
Sep 24

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

AgroBench is a reproducible benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel‑level crop time series. The data generation pipeline fuses USDA yield data with land cover masks, Sentinel‑2 and Sentinel‑1 imagery, climatic variables, and terrain information to produce multimodal sequences for individual crop pixels across the growing season. The benchmark includes over 13 million observations from 788,654 crop pixels, covering 5,107 county‑year combinations for five major U.S. crops from 2017 to 2024, and establishes a Leave‑One‑Year‑Out evaluation protocol with baseline machine learning results.

By Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat
arXiv Computation and Language
Sep 17

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA introduces a new panoptic grounded captioning framework that jointly generates detailed image captions and associates each phrase with precise pixel-level masks. The authors create PanoCaps, a human‑annotated benchmark with dense captions and near‑complete pixel coverage, and propose a phrase‑mask matching protocol with a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations, learns to select appropriate masks, and achieves state‑of‑the‑art grounding performance on PanoCaps and other pixel‑level tasks.

By Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid