arXiv Computer Vision

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

InterSing is a framework that generates realistic 3D head animations for duet singing by modeling the sparse, rhythm‑dependent interactions between performers. It introduces interaction logits—a weakly supervised, interpretable latent representation of cross‑performer engagement—and uses them to condition an interaction‑aware diffusion model driven by audio and interaction dynamics. The approach enables unified multi‑mode generation, producing coordinated behavior, independent motion, and smooth transitions, and it generalizes to multi‑singer performances with intuitive control over engagement.

arXiv Computer Vision
Sep 18

SalsaAgent: A multimodal embodied language model for interactive dance generation

SalsaAgent is a multimodal embodied language model that generates expressive, full‑body salsa follower motions in response to a human leader and music. The approach treats partner interaction as nonverbal token passing, extending a large language model’s vocabulary to include discrete motion, pairwise relation, and audio tokens. A two‑stage token‑to‑diffusion pipeline, combined with full‑body and pairwise‑relation tokenizers and alignment with automatically derived text descriptions of skeleton dynamics, yields improved motion quality, spatial coordination, and music‑partner synchrony compared to prior baselines.

By Payam Jome Yazdian, Zoe Stanley, Angelica Lim
Hugging Face Trending Papers
Jun 29

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models.

arXiv Computer Vision
Aug 27

InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

InteractGesture is a model‑agnostic, inference‑time method that enables fine‑grained spatial control of individual joints in continuous streaming co‑speech gesture generation. It guides diffusion sampler latent estimates through a differentiable RVQ‑VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. To address chunk‑wise dependency issues in streaming generation, the method introduces Progressive Chunk Guidance, a chunk‑window strategy that keeps an active set of editable chunk latents with staggered delays, allowing spatial constraints to propagate gradients backward across chunk boundaries and reducing boundary inconsistencies.

By Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli, Wenjing Zhang, Joanna Materzynska, Pu Wang, Vittorio Ferrari, Jie Shen
arXiv Computer Vision
Sep 11

Multi-Modal Controlled Coherent Motion Generation

The paper introduces MOCO, a diffusion-based framework that generates 3D avatar motions from concurrent multimodal inputs such as speech audio, text descriptions, and trajectory data. MOCO decouples motion generation by independently producing modality-specific motions at each denoising step and then assembling them according to spatial rules, iteratively refining the combined motion. This approach yields coherent, lifelike, and synchronized movements, outperforming existing baselines on a multimodal benchmark.

By Yifei Liu, Qiong Cao, Hongwei Yi, Huaiguang Jiang, Changxing Ding
arXiv Computer Vision
Sep 21

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.

By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei
arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv AI
Jul 16

Music-to-Dance Generation via Atomic Movements

arXiv:2607. 13978v1 Announce Type: cross Abstract: Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music.

By Xinhao Cai, Yixuan Sun, Minghang Zheng, Qingchao Chen, Xin Jin, Song-chun Zhu, Yang Liu