arXiv:2608.23636v1 Announce Type: new
Abstract: Small-object detection and instance segmentation remain challenging in orchard environments because of green-on-green similarity, occlusion, and limite...
By Ranjan Sapkota, Manoj Karkee
This study evaluates the instance‑segmentation performance of YOLOv11 and YOLOv8 on immature green apples in orchard settings. YOLO11n‑seg achieved the highest mask precision (0.831), while YOLO11m‑seg and YOLO11l‑seg excelled in non‑occluded and occluded fruitlet segmentation. YOLOv8n, however, outperformed the YOLO11 series in inference speed, reaching 3.3 ms compared to 4.8 ms for the fastest YOLO11 model.
By Ranjan Sapkota, Manoj Karkee
The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.
By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins
The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.
By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
arXiv:2608. 11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming.
By Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
This paper provides a detailed overview of the Ultralytics YOLO family from YOLOv5 to YOLO27, highlighting key architectural changes, benchmarking results, and deployment considerations. It discusses the evolution of each version—YOLO27’s scale‑adaptive dual architecture, YOLO26’s loss and optimization refinements, YOLO11’s efficiency focus, YOLOv8’s anchor‑free detection, and YOLOv5’s modular ecosystem—alongside performance metrics on COCO and latency on TensorRT. The review also surveys applications in robotics, agriculture, surveillance, and manufacturing, and outlines future challenges such as dense scene handling, CNN‑Transformer integration, and hardware‑aware optimization.
By Ranjan Sapkota, Manoj Karkee
Commercial greenhouse cucumber production is graded by fruit length, which drives harvest scheduling, labour allocation, and logistics. Manual measurement with thread or caliper is accurate but infeasible at commercial scale.
arXiv:2608.21454v1 Announce Type: new
Abstract: The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must...
By Subhankar Chattoraj, Sawon Pratiher, Samiran Das, Hubert Konik
arXiv:2609.23061v1 Announce Type: new
Abstract: Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds....
By Linduo Wei, Junjie Fan, Yijun Mai, Yong Qi
The paper evaluates the new YOLO26 architecture, which offers NMS-free end-to-end inference and is tailored for CPU-based edge devices, against three earlier Ultralytics models (YOLOv5u, YOLOv8, and YOLO11) in aquaculture fish mortality detection. Across nano, small, and medium scales, all models achieved similar detection accuracy on a full dataset, but differences emerged in data efficiency and deployment performance: YOLOv8 reached 90% mAP50 with only 400 images, while YOLO26 variants needed 1,000 images; YOLO26n was fastest on a Raspberry Pi 5 (7.51 FPS), whereas YOLOv5mu led on CPU-based hardware. The study concludes that architectural novelty alone does not dictate suitability for edge AI in aquaculture; training data size, target hardware, and inference needs must be jointly considered.
By Rakesh Ranjan, Gajanan S. Kothawade, Kata Sharrer, Scott Tsukuda, Christopher Good
arXiv:2409. 16808v3 Announce Type: replace-cross Abstract: Modern applications such as autonomous vehicles, intelligent surveillance, and smart city systems increasingly require object detection on resource-constrained edge devices.
By Daghash K. Alqahtani, Muhammad Aamir Cheema, Maria A. Rodriguez, Adel N. Toosi
arXiv:2606. 03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware.
By Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, Muhammet Esat Kalfaoglu