Hugging Face Trending Papers

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

arXiv AI
3d ago

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study. "whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."

By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
arXiv Computer Vision
Sep 1

RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

arXiv:2608.28693v1 Announce Type: cross Abstract: Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot inte...

By Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin, Tianze Wang, Zekun Qi, Liangliang Zhao, He Wang, Li Yi
arXiv AI
3d ago

GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots

GestAdapt is a framework that generates co‑speech gestures conditioned on a specified wrist workspace, allowing humanoid robots to adapt their motions to environmental constraints such as walls. The system learns from six co‑speech corpora using a shared motion representation and can be retargeted to different robot embodiments. Experiments show that GestAdapt’s motions stay close to real‑motion distributions, achieve higher quality scores than a no‑workspace baseline, and outperform other methods in real‑robot evaluations on the Reachy2 humanoid.

By Bosong Ding, Xianglin Zhang, Miao Xin, Murat Kirtay, Giacomo Spigler
arXiv AI
Sep 17

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

The paper introduces a pipeline that combines generated video and audio to produce force-aware manipulation trajectories for a Franka Panda robot. By using the loudness of contact sounds to shape a bounded, time-varying desired-force profile, the system can execute tasks that require precise contact forces, outperforming kinematic-only baselines. The approach also serves as a data generation engine for training closed-loop policies.

By Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
Hugging Face Trending Papers
Jul 8

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective.

arXiv Machine Learning
Sep 7

SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control

SCRIPT is a scalable diffusion policy that uses a Joint Action-State-Text Diffusion Transformer (JAST‑DiT) to jointly encode actions, physical states, and natural‑language instructions, enabling direct interaction between language semantics and control dynamics. The method employs a multi‑stage training framework, including supervised imitation pre‑training, a nonlinear history conditioning mechanism for stable autoregressive control, and a post‑training stage with Reinforcement Learning with Hybrid Rewards (RLHR) that injects learnable noise to improve motion quality and instruction following. Experiments on the 1200‑hour MotionMillion dataset show that SCRIPT outperforms prior state‑of‑the‑art methods across text alignment, motion quality, and physical realism, and its performance scales consistently with model size.

By Jingyan Zhang, Han Liang, Ruichi Zhang, Bin Li, Juze Zhang, Xin Chen, Jingya Wang, Lan Xu, Jingyi Yu