GROW$^2$: Grounding Which and Where for Robot Tool Use
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
The paper introduces a modular perception framework that uses vision‑language models (VLMs) to annotate object‑level regions from a single RGB‑D observation, then grounds these annotations with depth data to build an object‑centric representation. Experiments on 151 tabletop scenes demonstrate that this decomposition maintains strong semantic performance while significantly improving localization and depth estimation compared to direct VLM inference. The resulting representation is integrated into a task‑planning system for robotic manipulation.
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
arXiv:2606. 06761v1 Announce Type: cross Abstract: Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts.
arXiv:2506.11261v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and...
arXiv:2608. 19968v1 Announce Type: cross Abstract: Modern computer vision has enabled partial autonomy in robotic assembly manipulation.
V-Link is a method designed to enhance Vision‑Language‑Action (VLA) models by recovering visual representations during the transfer from vision‑language (VL) features to action (A) features. It introduces complementary Spatial and Semantic Query representations that are injected into Action DiT through asymmetric pathways, providing both semantic augmentation and dedicated geometric conditioning for action generation. Experiments on LIBERO, LIBERO‑Plus, RoboTwin 2.0, and real‑world AGIBOT A3 Ultra tasks show significant performance gains over the base GR00T N1.6 model.
arXiv:2605.03927v3 Announce Type: replace Abstract: Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle...
arXiv:2606. 10918v1 Announce Type: cross Abstract: The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios.
arXiv:2601.17885v2 Announce Type: replace-cross Abstract: Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint changes and scene variations. Exis...
arXiv:2607.11498v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting...
arXiv:2508.13073v3 Announce Type: replace-cross Abstract: Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditiona...
The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.
arXiv:2607. 08970v1 Announce Type: cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model.