arXiv Machine Learning By Jayant Ghadge, Soumyashree Kar, Surya S. Durbha

PhenoNEST: A Neuro-Symbolic Framework for Ontology-Aware Multimodal Plant Phenotyping and Trait Discovery

Read the original on arXiv Machine Learning →

arXiv:2607. 03245v1 Announce Type: new Abstract: High-throughput plant phenotyping generates valuable data that often remains trapped in unstructured text and isolated RGB images.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.

By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
arXiv Computer Vision
Sep 18

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.

By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
Hugging Face Trending Papers
Sep 17

AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding, supporting image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. It incorporates biologically specialized semantic representations with dense spatial grounding through biological‑semantic encoding, dense spatial representations, and pixel decoding. The authors also introduce AgriGround, a large‑scale dataset of over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask generation, and task‑oriented instruction synthesis to provide densely grounded supervision for agricultural vision‑language learning.

arXiv AI
Sep 24

Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning

The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.

By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles