arXiv Computer Vision

Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment

This study evaluates the instance‑segmentation performance of YOLOv11 and YOLOv8 on immature green apples in orchard settings. YOLO11n‑seg achieved the highest mask precision (0.831), while YOLO11m‑seg and YOLO11l‑seg excelled in non‑occluded and occluded fruitlet segmentation. YOLOv8n, however, outperformed the YOLO11 series in inference speed, reaching 3.3 ms compared to 4.8 ms for the fastest YOLO11 model.

arXiv Computer Vision
Sep 21

Optimizing YOLO27, YOLO26, YOLO11, and YOLOv8 for Fine-Grained Small-Object Detection and Segmentation in Complex Orchard Environments

The paper compares Ultralytics YOLO27, YOLO26, YOLO11, and YOLOv8 for detecting and segmenting small fruit parts in orchard settings. It evaluates five model scales across 30 experiments, finding that YOLO11s-960 and YOLO26s-960 achieve the best mask and box mAP scores while maintaining efficient parameter counts. The study also highlights the difficulty of peduncle detection and provides publicly available code and models for reproducibility.

By Ranjan Sapkota, Manoj Karkee
arXiv Computer Vision
Sep 11

Does YOLO26 Truly Offer Advantages Over Its Predecessors for Edge Deployment? A Benchmark Study in Aquaculture

The paper evaluates the new YOLO26 architecture, which offers NMS-free end-to-end inference and is tailored for CPU-based edge devices, against three earlier Ultralytics models (YOLOv5u, YOLOv8, and YOLO11) in aquaculture fish mortality detection. Across nano, small, and medium scales, all models achieved similar detection accuracy on a full dataset, but differences emerged in data efficiency and deployment performance: YOLOv8 reached 90% mAP50 with only 400 images, while YOLO26 variants needed 1,000 images; YOLO26n was fastest on a Raspberry Pi 5 (7.51 FPS), whereas YOLOv5mu led on CPU-based hardware. The study concludes that architectural novelty alone does not dictate suitability for edge AI in aquaculture; training data size, target hardware, and inference needs must be jointly considered.

By Rakesh Ranjan, Gajanan S. Kothawade, Kata Sharrer, Scott Tsukuda, Christopher Good
arXiv AI
Aug 12

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

arXiv:2608. 11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming.

By Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
arXiv Computer Vision
Sep 18

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.

By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins
arXiv Computer Vision
Aug 27

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
arXiv Computer Vision
Aug 25

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

The paper evaluates how well Vision Transformers (ViTs) can handle token merging techniques—specifically ToMe and Mutual Pair Merging—across wheat phenotyping tasks such as growth-stage classification, wheat-head detection, and wheat-organ segmentation. It benchmarks task quality, throughput, token count, and GPU memory usage, including tests on a Raspberry Pi 5. Results show that classification is highly tolerant to token merging, whereas detection and segmentation suffer due to factors like repeated instances, thin organs, dense boundaries, and runtime overhead, and that optimized attention backends can negate apparent speed gains.

By Simon Rav\'e, Pejman Rasti, David Rousseau
arXiv AI
Sep 17

Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper provides a detailed overview of the Ultralytics YOLO family from YOLOv5 to YOLO27, highlighting key architectural changes, benchmarking results, and deployment considerations. It discusses the evolution of each version—YOLO27’s scale‑adaptive dual architecture, YOLO26’s loss and optimization refinements, YOLO11’s efficiency focus, YOLOv8’s anchor‑free detection, and YOLOv5’s modular ecosystem—alongside performance metrics on COCO and latency on TensorRT. The review also surveys applications in robotics, agriculture, surveillance, and manufacturing, and outlines future challenges such as dense scene handling, CNN‑Transformer integration, and hardware‑aware optimization.

By Ranjan Sapkota, Manoj Karkee