arXiv Computer Vision By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

Read the original on arXiv Computer Vision →

The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 21

Optimizing YOLO27, YOLO26, YOLO11, and YOLOv8 for Fine-Grained Small-Object Detection and Segmentation in Complex Orchard Environments

The paper compares Ultralytics YOLO27, YOLO26, YOLO11, and YOLOv8 for detecting and segmenting small fruit parts in orchard settings. It evaluates five model scales across 30 experiments, finding that YOLO11s-960 and YOLO26s-960 achieve the best mask and box mAP scores while maintaining efficient parameter counts. The study also highlights the difficulty of peduncle detection and provides publicly available code and models for reproducibility.

By Ranjan Sapkota, Manoj Karkee
arXiv AI
Aug 12

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

arXiv:2608. 11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming.

By Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
arXiv Computer Vision
Aug 27

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
arXiv Computer Vision
Sep 4

DropClick: Semi-Automated One-Click Segmentation for Agricultural Robotic Data

DropClick is a click‑guided segmentation tool designed to ease the annotation of agricultural vision datasets. It uses single‑click inputs to generate pseudo‑labels, reducing the need for a click on every object and achieving high mIoU scores on SB20 and BUP20 datasets. The method also serves as a pseudo‑labelling approach that can train a Mask2Former model with significantly less user input while maintaining performance.

By Patrick Zimmer, Michael Halstead, Chris McCool
arXiv Computer Vision
Sep 14

RoMu4o: A Robotic Manipulation Unit For Orchard Operations Automating Proximal Hyperspectral Leaf Sensing

RoMu4o is a ground robot equipped with a 6‑DOF arm and a vision system that performs real‑time deep‑learning image processing and motion planning for proximal hyperspectral leaf sensing in orchards. The system uses robust perception and manipulation pipelines to identify leaf 3D structure, propose 6‑D poses, and generate collision‑free, constraint‑aware paths for precise leaf grasping and spectroscopy. In lab trials the robot achieved a 95 % success rate for 1‑LPB hyperspectral sampling, while field trials in a pistachio orchard reached 70 % success for autonomous leaf grasping and measurement. whyItMatters":"The system demonstrates a viable robotic solution to automate leaf‑level hyperspectral sensing, addressing labor shortages and enabling precise crop health monitoring in precision agriculture."

By Mehrad Mortazavi, David J. Cappelleri, Reza Ehsani