arXiv AI

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

arXiv Machine Learning
Jul 21

Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models

arXiv:2607. 16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometrically successful for the current skill can be brittle for downstream skills.

By Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias
arXiv Computer Vision
Aug 27

Training-Free Interaction-Aligned Visual Token Pruning for Efficient Embodied Manipulation

The paper introduces Interaction‑Aligned Pruning (IAprune), a training‑free method for visual token pruning in embodied manipulation tasks. IAprune jointly decides per‑frame budget and token selection, using semantic‑motion spatial agreement to choose between conservative and aggressive coverage, and applies geometric residual correction to focus on under‑represented boundaries. Experiments on four policies, three simulation benchmarks, and a real‑robot platform show that IAprune matches unpruned performance on LIBERO while achieving up to 1.54× speed‑up and 1.48× acceleration on a real robot.

By Jintao Cheng, Weibin Li, Haozhe Wang, Gang Wang, Yipu Zhang, Xiaoyu Tang, Jin Wu, Xieyuanli Chen, Yunhui Liu, Wei Zhang
arXiv AI
Sep 18

AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation

AntiGrounding is a visual action-selection framework that turns short robot trajectories into both executable motion plans and rendered prompts for vision‑language model evaluation. After filtering for feasibility, each trajectory is scored on safety, task alignment, efficiency, and physical plausibility using structured multi‑view visual question answering, and the best trajectories are refined and validated by a digital twin before real‑world execution. In eight real‑world manipulation tasks, the system achieved a 71.25% success rate with a single GPT‑6 Astra evaluator, outperforming baseline methods.

By Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu
arXiv AI
Sep 25

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

World Action Agent (WAA) is a multi‑agent framework that lets vision‑language models (VLMs) directly pilot robots by operating within a visual action workspace. The workspace provides automatically selected contact views, editable action rehearsals, and in‑view correction to refine decisions before low‑level execution. WAA learns procedural skills from expert videos and human teaching, and its interaction traces can train smaller VLMs, achieving state‑of‑the‑art success on LIBERO‑Pro and improving out‑of‑domain performance on robosuite and Qwen3.5‑9B.

By Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li
arXiv AI
Jun 3

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.

By Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
arXiv Computer Vision
Sep 18

INSPECT: Learning Robot View Selection from Assistant Use

INSPECT is a system that learns how a robot should choose its camera view during assembly inspection by observing a smart‑glasses assistant that answers part queries and guides the user. It uses techniques such as Presence‑Invariant TwinSwap for object evidence calibration, claim‑indexed supervision to separate evidence needs from camera changes, and object‑centered calibration to adapt view preferences to robot poses. In experiments on gearbox assemblies and angle‑grinder recordings, INSPECT outperforms other non‑oracle policies, improving view utility and decision accuracy.

By Di Wen, Kailun Yang, Wenhao Guo, Yitian Shi, Junwei Zheng, Yufan Chen, Ruiping Liu, Jiale Wei, Rania Rayyes, Kunyu Peng
arXiv Computer Vision
Sep 7

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

The paper proposes a neuro‑symbolic framework that augments vision‑language‑action (VLA) models with explicit task graphs and multimodal procedural memory to handle long‑horizon manipulation tasks. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory tracks the active step, completed actions, textual context, and relevant visual evidence. Human demonstrations provide spatial and temporal guidance via gaze or saliency cues, which are annotated in robot‑view teleoperation videos and used to fine‑tune VLA models. The approach is evaluated on workspace clearing and surgical‑instrument handling tasks, measuring object and destination selection, subtask completion, task progress, step‑order consistency, overall success, and procedural or execution mistakes.

By Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger
arXiv Machine Learning
Jun 16

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

arXiv:2511. 18960v4 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep.

By Lei Xiao, Jifeng Li, Juntao Gao, Feiyang Ye, Yan Jin, Jingjing Qian, Jing Zhang, Yong Wu, Xiaoyuan Yu