arXiv AI By Chongyu Bao, Haokai Yang, Yuhan Wang, Zhaochong An, Kunpeng Liu, Xiaolan Liu

Generative Embodied Multiple Behavior Control Systems for Human-like Agents

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 12

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

arXiv:2608. 10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication.

By Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Wang, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
Hugging Face Trending Papers
Jul 13

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences.

arXiv AI
Sep 17

A Comprehensive Review of Generative Physical Artificial Intelligence

The paper surveys Generative Physical Artificial Intelligence (GPAI), a field where large foundation models are integrated with physical robots. It introduces a taxonomy of five approaches—Robot Foundation Models, Vision‑Language Action models, Large Behavior Models, Diffusion Policy Models, and World Foundation Models—and discusses how they complement each other across domains such as autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems. The review highlights performance gains, data‑efficient learning, sim‑to‑real transfer, edge‑compatible architectures, and safety frameworks as key research directions.

By Satyam Gaba, Krutiksinh Rana, Siva Sai, Vinay Chamola, Dusit Niyato
arXiv Computer Vision
Sep 4

ProAct: Harnessing Streaming Motion Generation and Agentic Reasoning for Real-Time Embodied Social Interaction

ProAct is a dual‑system framework for real‑time embodied social interaction that separates a low‑latency Behavioral System, which streams multimodal interaction and generates continuous non‑verbal motion, from a slower Cognitive System that performs long‑horizon social reasoning and produces proactive intentions. The Cognitive System uses an efficient memory mechanism and a user‑motivation prediction module to decide when to intervene, while the Behavioral System translates these intentions into fluid motion via an intention‑conditioned streaming flow‑matching generator with a disentangled ControlNet branch. The framework is deployed on a physical humanoid robot and validated through real‑world user studies, motion‑generation benchmarks, and a new ProActBench benchmark for proactive trigger detection and restraint.

By Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng, Biao Jiang, Hanyu Ji, Libin Liu