arXiv AI

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

arXiv:2606. 12910v1 Announce Type: cross Abstract: For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time.

arXiv AI
Sep 4

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.

By Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv AI
3d ago

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

HiWE is a hierarchical world knowledge model that enables zero‑shot 3D path planning by linking visual grounding with language‑based planning through a point‑based interface. It uses PointVLM to map task‑relevant objects to image coordinates, lifts these predictions into a semantic 3D representation with depth data, and then a language planner (3DLLM) generates end‑effector waypoints and gripper commands. The system is evaluated on 14 simulated manipulation tasks and four physical‑robot tasks, with ablations on visual training data, spatial inputs, and grasp selection.

By Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu
arXiv AI
Jul 7

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.

By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
arXiv AI
Aug 19

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

The paper presents a perception-action pipeline for robotic pick‑and‑place in convenience stores, using annotation‑guided visual prompting to identify pickable objects and placement locations via bounding boxes. It replaces traditional step‑by‑step planning with Action Chunking with Transformers (ACT), an imitation learning algorithm that predicts chunked action sequences from human demonstrations. The system is evaluated on success rate and visual analysis of grasping behavior, showing improved grasp accuracy and adaptability in retail environments.

By Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae
arXiv AI
Jul 28

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.

By Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo