arXiv Computer Vision By Patrick Zimmer, Michael Halstead, Chris McCool

DropClick: Semi-Automated One-Click Segmentation for Agricultural Robotic Data

Read the original on arXiv Computer Vision →

DropClick is a click‑guided segmentation tool designed to ease the annotation of agricultural vision datasets. It uses single‑click inputs to generate pseudo‑labels, reducing the need for a click on every object and achieving high mIoU scores on SB20 and BUP20 datasets. The method also serves as a pseudo‑labelling approach that can train a Mask2Former model with significantly less user input while maintaining performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 18

SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.

By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
arXiv Machine Learning
Sep 11

Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.

By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan
arXiv Computer Vision
Sep 18

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.

By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins
arXiv AI
Sep 4

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.

By Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu