arXiv:2608.28693v1 Announce Type: cross
Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...
By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
GestAdapt is a framework that generates co‑speech gestures conditioned on a specified wrist workspace, allowing humanoid robots to adapt their motions to environmental constraints such as walls. The system learns from six co‑speech corpora using a shared motion representation and can be retargeted to different robot embodiments. Experiments show that GestAdapt’s motions stay close to real‑motion distributions, achieve higher quality scores than a no‑workspace baseline, and outperform other methods in real‑robot evaluations on the Reachy2 humanoid.
By Bosong Ding, Xianglin Zhang, Miao Xin, Murat Kirtay, Giacomo Spigler
arXiv:2606. 31158v1 Announce Type: cross Abstract: The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics.
By Snehasis Banerjee, Ranjan Dasgupta
The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.
By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
arXiv:2606. 19935v1 Announce Type: new Abstract: Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints.
By Zhangzhao Liang, Xiaofen Xing, Mingyue Yang, Wenlve Zhou, Xiangmin Xu
Auto-HSI is a system that creates personalized human‑swarm interaction interfaces on demand using large language models to automatically generate code from natural language descriptions and gesture demonstrations. The prototype enables untrained operators to control a swarm of 50 simulated robots with one‑ or two‑hand gestures, allowing teleoperation of motion, formation shape, and shape deformation. Experiments demonstrate the system’s gesture tracking, code generation, and live operation capabilities, including real‑time updates and deployment on real robots.
By Alessandro Nazzari, Nathan Cerisara, Dorian Tonnis, Raina Zakir, Lorenzo Labarile, Weixu Zhu, Marco Dorigo, Mary Katherine Heinrich