World Action Agent (WAA) is a multi‑agent framework that lets vision‑language models (VLMs) directly pilot robots by operating within a visual action workspace. The workspace provides automatically selected contact views, editable action rehearsals, and in‑view correction to refine decisions before low‑level execution. WAA learns procedural skills from expert videos and human teaching, and its interaction traces can train smaller VLMs, achieving state‑of‑the‑art success on LIBERO‑Pro and improving out‑of‑domain performance on robosuite and Qwen3.5‑9B.
By Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li
The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.
By Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao
arXiv:2601. 20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task-specific demonstrations and fine-tuning, and often generalize poorly under domain shift.
By Brian Y. Tsui, Alan Y. Fang, Tiffany J. Hwu
arXiv:2609.10522v1 Announce Type: cross
Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains cha...
By Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
arXiv:2607. 04162v1 Announce Type: cross Abstract: Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures.
By Iok Tong Lei, QianZhi Li, Ying Jie Yap, Yujie Zhang, Rui Zhong, Haichao Gui, Xiaolong Liu, Zhidong Deng
arXiv:2609.37810v1 Announce Type: cross
Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains chal...
By Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang