arXiv:2606.15848v2 Announce Type: replace
Abstract: 3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, an...
By Tingting Chen, Shaojun Wang, Huaye Zhang, Diqiong Jiang, Chenglizhao Chen
arXiv:2609.17422v1 Announce Type: new
Abstract: Audio-driven digital human generation plays an important role in virtual communication, immersive interaction, and media production. With the developme...
By Ziheng Yang, Yinfeng Yu, Yongming Li
The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.
By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv:2606. 28568v1 Announce Type: cross Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality.
By Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann, Rafael M. O. Cruz
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.
EmbedTalk introduces per‑Gaussian embeddings to drive speech‑driven facial deformations in real‑time talking head synthesis, replacing traditional tri‑plane encodings. This approach improves rendering quality, lip synchronisation, and motion consistency compared to prior 3D Gaussian Splatting methods while producing more compact models that run at 60+ FPS on a laptop GPU. The technique demonstrates competitive performance against state‑of‑the‑art generative models.
By Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya, David C. Hogg
arXiv:2609.18632v1 Announce Type: new
Abstract: Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusio...
By Qilin Wang, Mingyu Li, Hao Tang
Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still strugg...
arXiv:2606.00751v2 Announce Type: replace
Abstract: Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements. Still, its performance is fundamentally limited by...
By Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh
arXiv:2602. 07106v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for natural human-computer interaction.
By Haoyu Zhang, Zhipeng Li, Yiwen Guo, Tianshu Yu
arXiv:2606. 01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos.
By Zhicheng Zhang, Lei Wang, Yu Zhang, Yongsheng Gao
arXiv:2609.38019v1 Announce Type: new
Abstract: We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rat...
By Bangxun Tang