arXiv:2606. 19935v1 Announce Type: new Abstract: Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints.
By Zhangzhao Liang, Xiaofen Xing, Mingyue Yang, Wenlve Zhou, Xiangmin Xu
The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.
By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
arXiv:2606. 18747v1 Announce Type: cross Abstract: Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.
By Chris Lee, Flora Salim, Benjamin Tag, Francisco Cruz
arXiv:2609.00369v1 Announce Type: new
Abstract: Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging...
By Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
InteractGesture is a model‑agnostic, inference‑time method that enables fine‑grained spatial control of individual joints in continuous streaming co‑speech gesture generation. It guides diffusion sampler latent estimates through a differentiable RVQ‑VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. To address chunk‑wise dependency issues in streaming generation, the method introduces Progressive Chunk Guidance, a chunk‑window strategy that keeps an active set of editable chunk latents with staggered delays, allowing spatial constraints to propagate gradients backward across chunk boundaries and reducing boundary inconsistencies.
By Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli, Wenjing Zhang, Joanna Materzynska, Pu Wang, Vittorio Ferrari, Jie Shen
arXiv:2606. 31158v1 Announce Type: cross Abstract: The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics.
By Snehasis Banerjee, Ranjan Dasgupta
arXiv:2607. 14182v1 Announce Type: cross Abstract: Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies.
By J. M. A. Marcelo, M. Brienza, E. Bugli, L. Comito, D. Nardi, D. D. Bloisi, V. Suriani
DuoGesture is a co‑speech gesture generation model that separates gesture synthesis into a semantic stream and a beat stream, coordinated by a Semantic Variational Information Bottleneck that decides when semantic gestures override rhythmic motion. The semantic stream uses Motion‑Grounded Semantic Conditioning, replacing word embeddings with motion‑language representations to provide motion‑aligned semantic priors for rare gesture triggers. The beat stream is regularised by an Inertial Beat Prior, an anthropometry‑weighted arm‑chain module that reduces jitter and improves rhythmic consistency. Experiments show DuoGesture outperforms strong baselines and ablations confirm the complementary roles of semantic grounding, stochastic stream selection, and biomechanical regularisation.
By Ferdinand Paar, Lanmiao Liu, Asl{\i} \"Ozy\"urek, Serge Thill, Esam Ghaleb
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize au...
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).
By Bertram Taetz, Hugo Albuquerque Cosme da Silva, Gabriele Bleser-Taetz
arXiv:2607. 08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics.
By Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
arXiv:2606. 22726v2 Announce Type: replace Abstract: Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and expressive full-body dynamics.
By Seong Jong Yoo, Siyuan Peng, Felix Gu, Stratis Aloimonos, Cornelia Ferm\"uller