arXiv Computer Vision By Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng, Biao Jiang, Hanyu Ji, Libin Liu

ProAct: Harnessing Streaming Motion Generation and Agentic Reasoning for Real-Time Embodied Social Interaction

Read the original on arXiv Computer Vision →

ProAct is a dual‑system framework for real‑time embodied social interaction that separates a low‑latency Behavioral System, which streams multimodal interaction and generates continuous non‑verbal motion, from a slower Cognitive System that performs long‑horizon social reasoning and produces proactive intentions. The Cognitive System uses an efficient memory mechanism and a user‑motivation prediction module to decide when to intervene, while the Behavioral System translates these intentions into fluid motion via an intention‑conditioned streaming flow‑matching generator with a disentangled ControlNet branch. The framework is deployed on a physical humanoid robot and validated through real‑world user studies, motion‑generation benchmarks, and a new ProActBench benchmark for proactive trigger detection and restraint.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI introduces a streaming multimodal temporal modeling framework that enhances Vision‑Language‑Action models by adding temporal reasoning without extra parameters. It anchors each visual observation and language instruction pair as a temporal unit, using bidirectional attention for cross‑modal fusion and causal attention for autoregressive streaming inference. The method employs random‑interval streaming training to improve robustness and leverages the LLM backbone’s length extrapolation to inherit pretrained weights, achieving superior performance over pi0.5 on real‑robot and simulation tasks.

By Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya Yang, Xianzhe Fan, Junwei Luo, Junyi Li, Ruihua Han, Zhi Hou, Hengshuang Zhao
arXiv AI
Jul 7

iFLYTEK-Embodied-Omni Technical Report

arXiv:2607. 02542v1 Announce Type: new Abstract: General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons.

By Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang, Qingshan Xu, Chi Liu, Xin Nie, Wenjie Xu, Lin Gao, Zhiyuan Cheng, Mingxin Zhou, Jiajia Wu, Diyuan Liu, Jia Pan, Chao Ji