arXiv:2606. 28568v1 Announce Type: cross Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality.
By Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann, Rafael M. O. Cruz
InterSing is a framework that generates realistic 3D head animations for duet singing by modeling the sparse, rhythm‑dependent interactions between performers. It introduces interaction logits—a weakly supervised, interpretable latent representation of cross‑performer engagement—and uses them to condition an interaction‑aware diffusion model driven by audio and interaction dynamics. The approach enables unified multi‑mode generation, producing coordinated behavior, independent motion, and smooth transitions, and it generalizes to multi‑singer performances with intuitive control over engagement.
By Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu, Cheng Xu, Shengfeng He
arXiv:2606. 01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos.
By Zhicheng Zhang, Lei Wang, Yu Zhang, Yongsheng Gao
arXiv:2608. 11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing.
By Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models.
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
By Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji
Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.
By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv:2609.18632v1 Announce Type: new
Abstract: Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusio...
By Qilin Wang, Mingyu Li, Hao Tang
arXiv:2510. 02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content.
By Amir Dellali, Luca A. Lanzend\"orfer, Florian Gr\"otschla, Roger Wattenhofer
Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still strugg...
arXiv:2608.31106v1 Announce Type: new
Abstract: Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We...
By Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
arXiv:2606. 17126v1 Announce Type: cross Abstract: Singing style is a crucial aspect of a natural and expressive singing voice.
By Joon-Seung Choi, Dong-Min Byun, Seong-Whan Lee