arXiv:2606. 28568v1 Announce Type: cross Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality.
By Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann, Rafael M. O. Cruz
arXiv:2609.17422v1 Announce Type: new
Abstract: Audio-driven digital human generation plays an important role in virtual communication, immersive interaction, and media production. With the developme...
By Ziheng Yang, Yinfeng Yu, Yongming Li
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.
arXiv:2608. 05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact.
By Ao Fu, Yi Zhou
arXiv:2608. 15110v1 Announce Type: cross Abstract: Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization.
By Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li
arXiv:2602. 07106v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for natural human-computer interaction.
By Haoyu Zhang, Zhipeng Li, Yiwen Guo, Tianshu Yu
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail...
EmbedTalk introduces per‑Gaussian embeddings to drive speech‑driven facial deformations in real‑time talking head synthesis, replacing traditional tri‑plane encodings. This approach improves rendering quality, lip synchronisation, and motion consistency compared to prior 3D Gaussian Splatting methods while producing more compact models that run at 60+ FPS on a laptop GPU. The technique demonstrates competitive performance against state‑of‑the‑art generative models.
By Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya, David C. Hogg
arXiv:2503. 14295v3 Announce Type: replace-cross Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization.
By Baiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu, Zhen Lei
The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.
By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv:2606. 01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos.
By Zhicheng Zhang, Lei Wang, Yu Zhang, Yongsheng Gao
arXiv:2608.00663v2 Announce Type: replace
Abstract: Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods strug...
By Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun, Mingli Song, Jie Song