HarnessPAI introduces a model‑ and embodiment‑agnostic harness framework for Physical AI that treats code as an executable, evolvable interface organizing action primitives. The framework operates on two timescales: within a rollout it executes an open‑loop program, and across rollouts it evolves a closed‑loop program using feedback to refine the program and distill reusable skills. Across diverse robotic platforms, HarnessPAI outperforms pure action models and code‑as‑policy baselines, achieving significant gains on tasks such as LIBERO‑PRO and RoboCasa, and enabling efficient expert‑data collection for further fine‑tuning.
The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.
By Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao
EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.
By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
arXiv:2607. 11119v1 Announce Type: cross Abstract: Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control.
By Hengyuan Hu, Priya Sundaresan, Jensen Gao, Dorsa Sadigh
arXiv:2609.40306v1 Announce Type: cross
Abstract: Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and phys...
By Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang
arXiv:2607. 00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures.
By Runyu Lu, Yubo Wu, Ethan Kou, Letian Fu, Wenli Xiao, Ajay Mandlekar, Yinzhen Xu, Guanya Shi, Ken Goldberg, Ang Chen, Mosharaf Chowdhury, Yuke Zhu, Linxi "Jim" Fan, Guanzhi Wang