arXiv AI

CoMPAS3D: A Dataset and Benchmark for Interactive Motion

arXiv:2507. 19684v2 Announce Type: replace-cross Abstract: Socially interactive humanoid robots must engage with humans through their bodies, adapting in real time to a partner's movement, intent, and abilities.

arXiv Computer Vision
Sep 18

SalsaAgent: A multimodal embodied language model for interactive dance generation

SalsaAgent is a multimodal embodied language model that generates expressive, full‑body salsa follower motions in response to a human leader and music. The approach treats partner interaction as nonverbal token passing, extending a large language model’s vocabulary to include discrete motion, pairwise relation, and audio tokens. A two‑stage token‑to‑diffusion pipeline, combined with full‑body and pairwise‑relation tokenizers and alignment with automatically derived text descriptions of skeleton dynamics, yields improved motion quality, spatial coordination, and music‑partner synchrony compared to prior baselines.

By Payam Jome Yazdian, Zoe Stanley, Angelica Lim
Hugging Face Trending Papers
Jun 3

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granularity, undifferentiated difficulty, limited annotation quality, and pervasive answer ambiguity, leaving them unable to diagnose where current models fail.

arXiv Machine Learning
Jun 17

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

arXiv:2606. 17846v1 Announce Type: cross Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale.

By Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen
arXiv Computer Vision
Sep 7

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Motion-Omni is an end‑to‑end framework that jointly generates spoken dialogue and full‑body motion, producing speech, facial expressions, and hand, upper‑body, and lower‑body movements directly from the hidden states of a language model. The system requires joint training of the language model, speech generator, and motion generator to maintain audio‑motion alignment, and it is supervised using a scalable, model‑agnostic pipeline that pseudo‑labels 422,856 speech‑motion pairs. With a Qwen2.5‑7B‑Instruct backbone, Motion‑Omni‑Q7 achieves near‑cascade performance on motion metrics while being 5.4× faster, and it outperforms other non‑teacher cascades on beat correlation, diversity, and word error rate.

By Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo