arXiv AI

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.

arXiv AI
Sep 25

HarnessPAI: An Evolving Harness for Physical AI

HarnessPAI is a model‑ and embodiment‑agnostic framework that treats code as an executable, evolvable interface for Physical AI. It separates short‑term open‑loop program execution from long‑term closed‑loop evolution, using feedback to refine programs and distill reusable skills. Across various robots, HarnessPAI outperforms pure action models and code‑as‑policy baselines, achieving significant gains on tasks like LIBERO‑PRO and RoboCasa without retraining the underlying model.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan, Yang Li, Qing Li, Shangding Gu, Huichi Zhou, Shuqing Shi, Fei Ni, Shuo Lu, Weicheng Meng, Kang Li, Jin Wu, Kang Zhao, Shangmin Guo, Gen Li, Yongqiang Tang, Zhizhong Zhang, Yuan Xie, Heng Qu
Hugging Face Trending Papers
Sep 24

HarnessPAI: An Evolving Harness for Physical AI

HarnessPAI introduces a model‑ and embodiment‑agnostic harness framework for Physical AI that treats code as an executable, evolvable interface organizing action primitives. The framework operates on two timescales: within a rollout it executes an open‑loop program, and across rollouts it evolves a closed‑loop program using feedback to refine the program and distill reusable skills. Across diverse robotic platforms, HarnessPAI outperforms pure action models and code‑as‑policy baselines, achieving significant gains on tasks such as LIBERO‑PRO and RoboCasa, and enabling efficient expert‑data collection for further fine‑tuning.

arXiv AI
Jul 2

ASPIRE: Agentic /Skills Discovery for Robotics

arXiv:2607. 00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures.

By Runyu Lu, Yubo Wu, Ethan Kou, Letian Fu, Wenli Xiao, Ajay Mandlekar, Yinzhen Xu, Guanya Shi, Ken Goldberg, Ang Chen, Mosharaf Chowdhury, Yuke Zhu, Linxi "Jim" Fan, Guanzhi Wang
arXiv AI
Jun 19

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

arXiv:2606. 19980v1 Announce Type: new Abstract: Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence.

By Wenli Xiao, Jia Xie, Tonghe Zhang, Haotian Lin, Letian "Max" Fu, Haoru Xue, Jalen Lu, Yi Yang, Cunxi Dai, Zi Wang, Jimmy Wu, Guanzhi Wang, S. Shankar Sastry, Ken Goldberg, Linxi "Jim" Fan, Yuke Zhu, Guanya Shi
arXiv AI
Sep 25

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

World Action Agent (WAA) is a multi‑agent framework that lets vision‑language models (VLMs) directly pilot robots by operating within a visual action workspace. The workspace provides automatically selected contact views, editable action rehearsals, and in‑view correction to refine decisions before low‑level execution. WAA learns procedural skills from expert videos and human teaching, and its interaction traces can train smaller VLMs, achieving state‑of‑the‑art success on LIBERO‑Pro and improving out‑of‑domain performance on robosuite and Qwen3.5‑9B.

By Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li
arXiv Computer Vision
2d ago

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

arXiv:2610.01939v1 Announce Type: new Abstract: Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant obser...

By Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou
arXiv Computer Vision
Sep 22

An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

arXiv:2609.24170v1 Announce Type: new Abstract: Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency,...

By Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, Xiaoyu Huang, Liyang Li, Kailun Su, Weiyang Jin, Wenhao Chai, Haotian Liang, Zhiyang Dou, Yue Chen, Tianxing Chen
arXiv AI
Jul 7

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

arXiv:2607. 04927v1 Announce Type: cross Abstract: World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning.

By Jian Zhu, Jianjun Zhang, Taiyi Su, Tianbin Liu, Zhangyuan Wang, Kai Xie, Zitai Huang, Chong Ma, Youzhang He, Tianjian Wang, Hanyang Wang, Weihao Ding, Yi Xu
arXiv AI
Sep 18

Learning and Transferring Closed-Loop Robot Software

The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.

By So Kuroki, Yujin Tang
arXiv AI
Sep 2

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.

By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang