arXiv Computer Vision

AI-enabled Low-Cost 3D Maize Ear Morphometry Platform at Breeding Scale

arXiv Computer Vision
Sep 3

Automated Maize Ear Phenotyping Using 3D Reconstructions

The paper presents a fully automated pipeline that extracts maize ear traits—such as kernel count, row number, and kernel size—from 3D point clouds generated by a video-to-point-cloud platform. The method processes raw video through COLMAP and NeRF, isolates the ear, calibrates the point cloud, aligns it, unwraps it into a 2D image, and applies Cellpose‑SAM for instance segmentation, achieving high accuracy (kernel count R² = 0.921, MAPE = 10.33 %) on a held‑out dataset. The resulting multi‑trait dataset, with genotype identities, is ready for phenotype‑to‑genotype association studies.

By Ritwesh A. Kumar, Som Tripathi, Peja Matthews, Srikar Reddy, Talukder Zaki Jubery, Patrick Schnable, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
arXiv Machine Learning
Aug 10

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

arXiv:2608. 06404v1 Announce Type: cross Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response.

By Junxiong Zhou, Xuechen Li, Chonghao Qiu, Lang Qiao, Xiaowei Jia, Qi Yang, Chishan Zhang, Leikun Yin, Nanshan You, Vipin Kumar, David Mulla, Ce Yang, Zhenong Jin, Licheng Liu
Hugging Face Trending Papers
Jul 2

The Turning Point of 3D Plant Phenotyping: 3D Foundation Models Enable Minute-to-Second Cross-Crop Reconstruction and Beyond

3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion.

arXiv AI
Sep 7

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

The paper presents a unified compression framework for Vision Transformers aimed at on‑device plant disease detection in resource‑constrained agricultural settings. It combines Hessian‑Balanced Adaptive Block Pruning, quantization, and attention‑based knowledge distillation, evaluating each component separately before integrating the best performers into a deployment pipeline. On a chilli disease dataset, the compressed models achieve accuracy comparable to the FP32 baseline while reducing model size by 74‑98 %, and the full pipeline attains a 54.5× size reduction to 6.01 MB with 95.13 % accuracy.

By Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi, Sudheesh Manalil, Arnab Raha, Amitava Mukherjee, Parthasarathy Seethapathy, G. Gopakumar
arXiv Computer Vision
Aug 25

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

The paper evaluates how well Vision Transformers (ViTs) can handle token merging techniques—specifically ToMe and Mutual Pair Merging—across wheat phenotyping tasks such as growth-stage classification, wheat-head detection, and wheat-organ segmentation. It benchmarks task quality, throughput, token count, and GPU memory usage, including tests on a Raspberry Pi 5. Results show that classification is highly tolerant to token merging, whereas detection and segmentation suffer due to factors like repeated instances, thin organs, dense boundaries, and runtime overhead, and that optimized attention backends can negate apparent speed gains.

By Simon Rav\'e, Pejman Rasti, David Rousseau
arXiv Computer Vision
Sep 16

Evaluating Mesh Reconstruction Methods for Crop Phenotyping

The paper evaluates seven 3D mesh reconstruction pipelines for crop phenotyping, assessing their fidelity and consistency both qualitatively and quantitatively. Results indicate that the GGGS, PGSR, and 2DGS pipelines produce the most accurate and visually pleasing meshes, with GGGS outperforming the next best (2DGS) by about 27% across five metrics: User ratings, Chamfer distance, LPIPS, PSNR, and SSIM.

By Karanvir Singh, Theo Morales, Binh-Son Hua, Mukesh Saini
arXiv Computer Vision
Sep 21

Optimizing YOLO27, YOLO26, YOLO11, and YOLOv8 for Fine-Grained Small-Object Detection and Segmentation in Complex Orchard Environments

The paper compares Ultralytics YOLO27, YOLO26, YOLO11, and YOLOv8 for detecting and segmenting small fruit parts in orchard settings. It evaluates five model scales across 30 experiments, finding that YOLO11s-960 and YOLO26s-960 achieve the best mask and box mAP scores while maintaining efficient parameter counts. The study also highlights the difficulty of peduncle detection and provides publicly available code and models for reproducibility.

By Ranjan Sapkota, Manoj Karkee
arXiv Computer Vision
Aug 31

Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait Extraction

The paper introduces SynthCrop4D, a synthetic dataset of temporally evolving plant point clouds that includes controllable noise, occlusion, and complete geometry for benchmarking reconstruction methods. It proposes a two‑stage pipeline combining spatial denoising with an Adaptive Temporal PoinTr model to recover missing regions from self‑occlusion, achieving significant improvements in reconstruction quality on both SynthCrop4D and the real Pheno4D dataset. The completed point clouds are further used to extract phenotypic traits such as plant height, canopy width, and convex hull volume, demonstrating the pipeline’s utility for high‑throughput crop phenotyping.

By Mrudul Mittal, Soumyashree Kar