arXiv Computer Vision
Aug 27

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

The paper introduces a lightweight multimodal vision‑language framework based on TinyCLIP for fine‑grained classification of early‑stage apple fruitlet anatomy (calyx, fruitlet body, peduncle) in orchard images. Using a dataset of 600 high‑resolution RGB images, the model employs domain‑specific language prompts and a sliding‑window inference strategy to produce interpretable heatmaps for whole‑image localization. Achieving macro‑F1 of 0.93 on an NVIDIA T4 GPU and maintaining accuracy after INT8 quantization, the system is optimized for edge deployment on NVIDIA Jetson hardware with model sizes around 127‑137 MB and millisecond‑level inference.

By Ranjan Sapkota, William Bu, Chen Chen, Yunjun Xu, Manoj Karkee
arXiv Computer Vision
Sep 3

MultiGraspNet: A Multitask 3D Vision Model for Multi-gripper Robotic Grasping

MultiGraspNet is a multitask 3D vision model that simultaneously predicts feasible poses for both parallel and vacuum grippers, allowing a single robot to handle multiple end effectors. Trained on the aligned GraspNet-1Billion and SuctionNet-1Billion datasets, it generates graspability masks that quantify the suitability of each scene point for successful grasps. With only 15.75 M parameters, the model achieves fast inference on a single GPU and demonstrates competitive performance against single-task models while reducing computational cost, as shown in extensive experiments and real‑world tests on a single‑arm multi‑gripper setup.

By Stephany Ortuno-Chanelo, Paolo Rabino, Enrico Civitelli, Tatiana Tommasi, Raffaello Camoriano
arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem