The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.
By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
arXiv:2609.00369v1 Announce Type: new
Abstract: Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging...
By Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize au...
arXiv:2604.09057v3 Announce Type: replace
Abstract: Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible...
By Junchao Liao, Zhenghao Zhang, Xiangyu Meng, Litao Li, Ziying Zhang, Siyu Zhu, Long Qin, Weizhi Wang
arXiv:2608.28693v1 Announce Type: cross
Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...
By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
The paper introduces MOCO, a diffusion-based framework that generates 3D avatar motions from concurrent multimodal inputs such as speech audio, text descriptions, and trajectory data. MOCO decouples motion generation by independently producing modality-specific motions at each denoising step and then assembling them according to spatial rules, iteratively refining the combined motion. This approach yields coherent, lifelike, and synchronized movements, outperforming existing baselines on a multimodal benchmark.
By Yifei Liu, Qiong Cao, Hongwei Yi, Huaiguang Jiang, Changxing Ding