arXiv Computer Vision

VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

The paper introduces a modular perception framework that uses vision‑language models (VLMs) to annotate object‑level regions from a single RGB‑D observation, then grounds these annotations with depth data to build an object‑centric representation. Experiments on 151 tabletop scenes demonstrate that this decomposition maintains strong semantic performance while significantly improving localization and depth estimation compared to direct VLM inference. The resulting representation is integrated into a task‑planning system for robotic manipulation.

arXiv AI
Jun 8

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

arXiv:2606. 06761v1 Announce Type: cross Abstract: Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts.

By Jiyun Jang, Yujin Sung, Woosung Joung, Daewon Chae, Sangwon Lee, Sohwi Kim, Jinkyu Kim, Jungbeom Lee
arXiv Computer Vision
Aug 27

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

V-Link is a method designed to enhance Vision‑Language‑Action (VLA) models by recovering visual representations during the transfer from vision‑language (VL) features to action (A) features. It introduces complementary Spatial and Semantic Query representations that are injected into Action DiT through asymmetric pathways, providing both semantic augmentation and dedicated geometric conditioning for action generation. Experiments on LIBERO, LIBERO‑Plus, RoboTwin 2.0, and real‑world AGIBOT A3 Ultra tasks show significant performance gains over the base GR00T N1.6 model.

By Yehao Lu, Jiarui Yang, Yuning Su, Yufeng Xie, Yu Zhong, Yazhou Zhang, Haiyu Lan, Kaixiang Lu, Peiwen Lin, Chuang Wang, Zequn Qin, Enyu Li, Xi Li
Hugging Face Trending Papers
Sep 8

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.