arXiv AI

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

arXiv:2606. 19935v1 Announce Type: new Abstract: Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints.

arXiv AI
3d ago

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study. "whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."

By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
arXiv Computer Vision
Sep 1

RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

arXiv:2608.28693v1 Announce Type: cross Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...

By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
arXiv AI
Jul 1

A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

arXiv:2509. 15443v2 Announce Type: replace-cross Abstract: Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing abundant, large-scale human motion collections.

By Xingyu Chen, Hanyu Wu, Sikai Wu, Mingliang Zhou, Diyun Xiang, Haodong Zhang, Yangchen Zhou, Yukang Gao, Yi Gu, Renjing Xu
arXiv AI
3d ago

GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots

GestAdapt is a framework that generates co‑speech gestures conditioned on a specified wrist workspace, allowing humanoid robots to adapt their motions to environmental constraints such as walls. The system learns from six co‑speech corpora using a shared motion representation and can be retargeted to different robot embodiments. Experiments show that GestAdapt’s motions stay close to real‑motion distributions, achieve higher quality scores than a no‑workspace baseline, and outperform other methods in real‑robot evaluations on the Reachy2 humanoid.

By Bosong Ding, Xianglin Zhang, Miao Xin, Murat Kirtay, Giacomo Spigler
Hugging Face Trending Papers
Sep 3

BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI

The paper presents BRIDGE, an open‑source 88 cm tall humanoid robot designed through a data‑driven morphology‑control co‑design framework that aligns robot shape with human‑like movement. It introduces a new metric combining kinematic retargeting fidelity and dynamic tracking performance to evaluate morphological fidelity, achieving state‑of‑the‑art results against baseline humanoids. The released platform, along with its control policy, demonstrates superior human motion capture, robust balance, and dynamic maneuvers, with supporting videos and code available online.

arXiv AI
Sep 4

BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI

The paper introduces BRIDGE, an open‑source 88 cm tall humanoid robot designed through a data‑driven morphology‑control co‑design framework that optimizes the robot’s body shape for human‑like movement. A new metric combining kinematic retargeting fidelity and dynamic tracking performance is proposed to evaluate morphological fidelity, and the framework achieves state‑of‑the‑art results compared to existing humanoids such as Bumi, K1, and Toddlerbot. The resulting platform, released with its control policy and supporting materials, demonstrates superior fidelity in capturing human motion, robust balance, and highly dynamic maneuvers.

By Jianren Wang, Letian Qian, Zikai Wang, Weiwei Wu, Junjie Zong, Abhinav Gupta, Deepak Pathak
arXiv AI
Sep 16

SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating

SafeFlow is a real‑time, text‑driven humanoid control framework that blends physics‑guided motion generation with a three‑stage safety gate. It uses Physics‑Guided Rectified Flow Matching in a VAE latent space to produce physically executable trajectories, accelerates sampling with Reflow, and filters unsafe outputs via semantic OOD detection, directional sensitivity checks, and hard kinematic constraints before handing them to a motion‑tracking controller. Experiments on the Unitree G1 show that SafeFlow achieves higher success rates, better physical compliance, and faster inference than diffusion‑ and retargeting‑based baselines while maintaining motion diversity.

By Hanbyel Cho, Sang-Hun Kim, Jeonguk Kang, Donghan Koo