arXiv AI

Music-to-Dance Generation via Atomic Movements

arXiv:2607. 13978v1 Announce Type: cross Abstract: Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music.

Hugging Face Trending Papers
Jun 29

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models.

arXiv Computer Vision
Sep 18

SalsaAgent: A multimodal embodied language model for interactive dance generation

SalsaAgent is a multimodal embodied language model that generates expressive, full‑body salsa follower motions in response to a human leader and music. The approach treats partner interaction as nonverbal token passing, extending a large language model’s vocabulary to include discrete motion, pairwise relation, and audio tokens. A two‑stage token‑to‑diffusion pipeline, combined with full‑body and pairwise‑relation tokenizers and alignment with automatically derived text descriptions of skeleton dynamics, yields improved motion quality, spatial coordination, and music‑partner synchrony compared to prior baselines.

By Payam Jome Yazdian, Zoe Stanley, Angelica Lim
arXiv Computer Vision
Sep 7

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

InterSing is a framework that generates realistic 3D head animations for duet singing by modeling the sparse, rhythm‑dependent interactions between performers. It introduces interaction logits—a weakly supervised, interpretable latent representation of cross‑performer engagement—and uses them to condition an interaction‑aware diffusion model driven by audio and interaction dynamics. The approach enables unified multi‑mode generation, producing coordinated behavior, independent motion, and smooth transitions, and it generalizes to multi‑singer performances with intuitive control over engagement.

By Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu, Cheng Xu, Shengfeng He