ProAct: Harnessing Streaming Motion Generation and Agentic Reasoning for Real-Time Embodied Social Interaction
Read the original on arXiv Computer Vision →ProAct is a dual‑system framework for real‑time embodied social interaction that separates a low‑latency Behavioral System, which streams multimodal interaction and generates continuous non‑verbal motion, from a slower Cognitive System that performs long‑horizon social reasoning and produces proactive intentions. The Cognitive System uses an efficient memory mechanism and a user‑motivation prediction module to decide when to intervene, while the Behavioral System translates these intentions into fluid motion via an intention‑conditioned streaming flow‑matching generator with a disentangled ControlNet branch. The framework is deployed on a physical humanoid robot and validated through real‑world user studies, motion‑generation benchmarks, and a new ProActBench benchmark for proactive trigger detection and restraint.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.