arXiv:2608.28693v1 Announce Type: cross
Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...
By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
arXiv:2606. 18747v1 Announce Type: cross Abstract: Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.
By Chris Lee, Flora Salim, Benjamin Tag, Francisco Cruz
arXiv:2607. 14182v1 Announce Type: cross Abstract: Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies.
By J. M. A. Marcelo, M. Brienza, E. Bugli, L. Comito, D. Nardi, D. D. Bloisi, V. Suriani
arXiv:2606. 19914v1 Announce Type: cross Abstract: Art has long stood as a pivotal expression of human creativity.
By Xuetao Li, Wenke Huang, Mang Ye, Zijian Liu, Jinhua Xie, Jifeng Xuan, Miao Li
ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study.
"whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."
By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information,...
The paper introduces a pipeline that combines generated video and audio to produce force-aware manipulation trajectories for a Franka Panda robot. By using the loudness of contact sounds to shape a bounded, time-varying desired-force profile, the system can execute tasks that require precise contact forces, outperforming kinematic-only baselines. The approach also serves as a data generation engine for training closed-loop policies.
By Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
arXiv:2603.16086v2 Announce Type: replace-cross
Abstract: While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts...
By Chang Nie, Tianchen Deng, Guangming Wang, Zhe Liu, Hesheng Wang
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly...
The paper introduces an LLM-based Conversational AI Knowledge Assistant for the Raspberry‑Pi‑powered 13‑Axis MyBuddy humanoid robot. It combines large language model-driven language understanding, real‑time speech recognition, internet‑based knowledge retrieval (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to support intelligent, multi‑turn conversations and emotional‑support interactions. This system aims to overcome the limitations of traditional rule‑based dialogue systems in humanoid robots.
By Hanxiao Chen
arXiv:2608. 16503v1 Announce Type: cross Abstract: Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness.
By Cong Zhao, Shuai Tian, Xu Zhang, Baocheng Ni, Xinguo Song, Xueying Sun, Shu Jiang, Shouchang Yang, Bo Tang, Jin Deng, Ge Zhu, YongCheng Wang, Jin Xu, Ri Yang
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA...