ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study.
"whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."
By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
arXiv:2606. 19935v1 Announce Type: new Abstract: Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints.
By Zhangzhao Liang, Xiaofen Xing, Mingyue Yang, Wenlve Zhou, Xiangmin Xu
arXiv:2608.28693v1 Announce Type: cross
Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...
By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
arXiv:2607. 14182v1 Announce Type: cross Abstract: Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies.
By J. M. A. Marcelo, M. Brienza, E. Bugli, L. Comito, D. Nardi, D. D. Bloisi, V. Suriani
Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still strugg...
GestAdapt is a framework that generates co‑speech gestures conditioned on a specified wrist workspace, allowing humanoid robots to adapt their motions to environmental constraints such as walls. The system learns from six co‑speech corpora using a shared motion representation and can be retargeted to different robot embodiments. Experiments show that GestAdapt’s motions stay close to real‑motion distributions, achieve higher quality scores than a no‑workspace baseline, and outperform other methods in real‑robot evaluations on the Reachy2 humanoid.
By Bosong Ding, Xianglin Zhang, Miao Xin, Murat Kirtay, Giacomo Spigler
arXiv:2609.18632v1 Announce Type: new
Abstract: Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusio...
By Qilin Wang, Mingyu Li, Hao Tang
The paper introduces a pipeline that combines generated video and audio to produce force-aware manipulation trajectories for a Franka Panda robot. By using the loudness of contact sounds to shape a bounded, time-varying desired-force profile, the system can execute tasks that require precise contact forces, outperforming kinematic-only baselines. The approach also serves as a data generation engine for training closed-loop policies.
By Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information,...
Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective.
arXiv:2608. 16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions.
By Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han
SCRIPT is a scalable diffusion policy that uses a Joint Action-State-Text Diffusion Transformer (JAST‑DiT) to jointly encode actions, physical states, and natural‑language instructions, enabling direct interaction between language semantics and control dynamics. The method employs a multi‑stage training framework, including supervised imitation pre‑training, a nonlinear history conditioning mechanism for stable autoregressive control, and a post‑training stage with Reinforcement Learning with Hybrid Rewards (RLHR) that injects learnable noise to improve motion quality and instruction following. Experiments on the 1200‑hour MotionMillion dataset show that SCRIPT outperforms prior state‑of‑the‑art methods across text alignment, motion quality, and physical realism, and its performance scales consistently with model size.
By Jingyan Zhang, Han Liang, Ruichi Zhang, Bin Li, Juze Zhang, Xin Chen, Jingya Wang, Lan Xu, Jingyi Yu