arXiv AI

Active sensing to characterize the heterogeneity of plant stress

The paper introduces an autonomous robotic platform that performs targeted chlorophyll fluorescence measurements on plant leaves. It integrates 3D plant reconstruction, geometric analysis, and motion planning to identify suitable leaf surfaces and generate collision‑free trajectories for a robotic manipulator. This system enables automated, repeatable, and spatially resolved physiological measurements that extend beyond passive imaging.

arXiv Computer Vision
Sep 14

RoMu4o: A Robotic Manipulation Unit For Orchard Operations Automating Proximal Hyperspectral Leaf Sensing

RoMu4o is a ground robot equipped with a 6‑DOF arm and a vision system that performs real‑time deep‑learning image processing and motion planning for proximal hyperspectral leaf sensing in orchards. The system uses robust perception and manipulation pipelines to identify leaf 3D structure, propose 6‑D poses, and generate collision‑free, constraint‑aware paths for precise leaf grasping and spectroscopy. In lab trials the robot achieved a 95 % success rate for 1‑LPB hyperspectral sampling, while field trials in a pistachio orchard reached 70 % success for autonomous leaf grasping and measurement. whyItMatters":"The system demonstrates a viable robotic solution to automate leaf‑level hyperspectral sensing, addressing labor shortages and enabling precise crop health monitoring in precision agriculture."

By Mehrad Mortazavi, David J. Cappelleri, Reza Ehsani
Hugging Face Trending Papers
Aug 4

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis

Plant root phenotyping is fundamental to understanding below-ground structures, optimizing crop management, and improving agricultural sustainability. This paper presents a multimodal robotic AI framework that integrates 3D skeleton extraction with language-guided reasoning for interpretable and data-efficient root analysis.

Hugging Face Trending Papers
Jul 2

The Turning Point of 3D Plant Phenotyping: 3D Foundation Models Enable Minute-to-Second Cross-Crop Reconstruction and Beyond

3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion.

arXiv Machine Learning
Jul 7

Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance

arXiv:2603. 07866v3 Announce Type: replace-cross Abstract: Offshore inspection and maintenance have increasingly been using legged robots for routine sensing, yet many useful interventions still require physical interaction with tools, containers, and task-relevant objects.

By Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto, Ranulfo Bezerra, Gustavo J. G. Lahr, Ricardo V. Godoy, Marcelo Becker
arXiv Computer Vision
Aug 27

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
Hugging Face Trending Papers
Jun 9

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes.

Hugging Face Trending Papers
Aug 18

Scale Matters: Adaptive Granularity Selection for Cross-Species 3D Plant Organ Segmentation

The paper introduces AGS-PlantSeg, a few‑shot 3D plant organ segmentation method that uses the frozen Utonia foundation model and Adaptive Granularity Selection (AGS) to dynamically choose optimal spatial granularity for each plant. By extracting tailored geometric features for a lightweight MLP head, AGS-PlantSeg achieves superior cross‑species generalization, reaching an average mIoU of 88.9% and outperforming fixed‑granularity baselines by 2.5 points across PLANesT‑3D, Pheno4D, and Crops3D datasets. The approach requires minimal annotated data yet competes with fully supervised, plant‑specific architectures.