Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.
arXiv:2606. 29148v1 Announce Type: cross Abstract: Developing controllers capable of completing a wide range of tasks in a natural and life-like manner is a key challenge in enabling practical applications of physics-based character animation.
By Yi Shi, Yifeng Jiang, Chen Tessler, Xue Bin Peng
The paper surveys Generative Physical Artificial Intelligence (GPAI), a field where large foundation models are integrated with physical robots. It introduces a taxonomy of five approaches—Robot Foundation Models, Vision‑Language Action models, Large Behavior Models, Diffusion Policy Models, and World Foundation Models—and discusses how they complement each other across domains such as autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems. The review highlights performance gains, data‑efficient learning, sim‑to‑real transfer, edge‑compatible architectures, and safety frameworks as key research directions.
By Satyam Gaba, Krutiksinh Rana, Siva Sai, Vinay Chamola, Dusit Niyato
arXiv:2509. 06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robotics.
By Yifei Ren, Edward Johns
The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI...
CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.
By Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology, China), Yukang Ding (Taobao, Alibaba Group, China), Hui Li (Harbin Institute of Technology, China), Wangmeng Zuo (Harbin Institute of Technology, China)