arXiv Computer Vision

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

arXiv Computer Vision
4d ago

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

The paper evaluates how well Vision Transformers (ViTs) can handle token merging techniques—specifically ToMe and Mutual Pair Merging—across wheat phenotyping tasks such as growth-stage classification, wheat-head detection, and wheat-organ segmentation. It benchmarks task quality, throughput, token count, and GPU memory usage, including tests on a Raspberry Pi 5. Results show that classification is highly tolerant to token merging, whereas detection and segmentation suffer due to factors like repeated instances, thin organs, dense boundaries, and runtime overhead, and that optimized attention backends can negate apparent speed gains.

By Simon Rav\'e, Pejman Rasti, David Rousseau
arXiv AI
Jun 2

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.

By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
Hugging Face Trending Papers
Aug 4

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis

Plant root phenotyping is fundamental to understanding below-ground structures, optimizing crop management, and improving agricultural sustainability. This paper presents a multimodal robotic AI framework that integrates 3D skeleton extraction with language-guided reasoning for interpretable and data-efficient root analysis.

arXiv Machine Learning
Aug 4

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

arXiv:2608. 01202v1 Announce Type: cross Abstract: Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management.

By Ahmed Baha Ben Jmaa, Faten Chaieb, Anna Fabija\'nska
arXiv AI
Jun 26

Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

arXiv:2606. 26151v1 Announce Type: cross Abstract: While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge.

By Th\'eo Biardeau (XLIM-ASALI, UFR SFA), Anne-Sophie Capelle-Laiz\'e (UP, XLIM-ASALI, XLIM-ASALI), Salwan Alwan (UFR SFA), David Helbert (UFR SFA)
arXiv AI
Aug 3

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.

By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv Computer Vision
2d ago

Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

The paper introduces the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), which fuses decision-level outputs from EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by multimodal large language models Gemma 4 E4B and Qwen3.5 4B to produce explainable plant disease diagnoses. Evaluated on 14,364 images from PlantDoc and two Cornell robotic field datasets, the framework achieves up to 99.3% accuracy, with Gemma improving PlantDoc accuracy from 63.9% to 68.5% and demonstrating low critical‑risk error. The results highlight the potential of MLLM arbitration for reliable, explainable agricultural AI under real‑world field conditions.

By Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee