arXiv Computer Vision

Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition

arXiv Machine Learning
Aug 4

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

arXiv:2608. 01202v1 Announce Type: cross Abstract: Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management.

By Ahmed Baha Ben Jmaa, Faten Chaieb, Anna Fabija\'nska
arXiv Computer Vision
5d ago

A Dataset-Centric Benchmark of Deep Learning Methods for Grape Leaf Disease Classification and Detection

The paper introduces a dataset‑centric benchmark for deep learning approaches to grape leaf disease classification and detection. It evaluates publicly available datasets on disease categories, annotations, acquisition conditions, and class distributions, and tests representative models across image‑level classification, region‑level classification, and object detection. Results reveal high accuracy on controlled datasets but significant performance drops on heterogeneous, real‑world data, especially in cross‑dataset transfer and object detection tasks.

By Petar Canoski, Vlatko Spasev, Ivica Dimitrovski, Ivan Kitanovski, Petre Lameski
arXiv AI
3d ago

STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification

STA‑Net is a lightweight neural network designed for plant disease classification on edge devices. It combines a training‑free neural architecture search (DeepMAD) to build an efficient backbone with a novel Shape‑Texture Attention Module (STAM) that separates shape and texture processing using deformable convolutions and a Gabor filter bank. On the CCMT plant disease dataset, STA‑Net achieved 89.00% accuracy and 88.96% F1 score with only 401K parameters and 51.1M FLOPs.

By Zongsen Qiu, Jianjun Wang, Yue Zhou, Zibo Zhou, Rui Chen
arXiv Computer Vision
2d ago

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
Hugging Face Trending Papers
Aug 2

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management. Accurate and timely FRP can be achieved using machine/deep learning-based hyperspectral image classification techniques.

arXiv AI
Jun 2

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.

By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
arXiv Computer Vision
4d ago

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

The paper evaluates how well Vision Transformers (ViTs) can handle token merging techniques—specifically ToMe and Mutual Pair Merging—across wheat phenotyping tasks such as growth-stage classification, wheat-head detection, and wheat-organ segmentation. It benchmarks task quality, throughput, token count, and GPU memory usage, including tests on a Raspberry Pi 5. Results show that classification is highly tolerant to token merging, whereas detection and segmentation suffer due to factors like repeated instances, thin organs, dense boundaries, and runtime overhead, and that optimized attention backends can negate apparent speed gains.

By Simon Rav\'e, Pejman Rasti, David Rousseau
arXiv AI
Aug 5

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.

By Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
arXiv Machine Learning
5d ago

On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift

arXiv:2608.21254v1 Announce Type: cross Abstract: Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reduci...

By Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez, Shekhar Borah, Athresh Karanam, Erik Blasch, Prabha Sundaravadivel, Sriraam Natarajan
arXiv AI
Jul 7

PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification

arXiv:2405. 07332v2 Announce Type: cross Abstract: Numerous applications have resulted from the automation of agricultural disease segmentation using deep learning techniques.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Mohammad Shafiul Alam, Ahmed Al Wase, Md. Rabius Sani, Khan Md Hasib
arXiv Computer Vision
2d ago

MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection

MIMONet is a saliency detection model that uses multi‑scale inputs and outputs to better handle objects of varying sizes. It processes three differently sized images through separate encoder branches that exchange information, allowing each branch to learn size‑variation knowledge from the others. A Multi‑scale Perception module further refines features, and a Joint Saliency Loss ensures consistent, well‑preserved boundaries across the multiple saliency maps produced.

By Zhaojian Yao, Wei Gao, Tiesong Zhao, Hui Yuan, Sam Kwong