The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.
By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv:2609.18632v1 Announce Type: new
Abstract: Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusio...
By Qilin Wang, Mingyu Li, Hao Tang
Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still strugg...
arXiv:2608. 05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact.
By Ao Fu, Yi Zhou
The paper introduces a geometric adaptation framework for cross‑speaker acoustic‑to‑articulatory inversion that leverages anatomical landmarks on vertebrae and dental structures. By applying an affine transformation followed by a thin‑plate spline deformation, the method maps predicted vocal‑tract contours from a fixed model to unseen speakers without retraining. Experiments on a single‑speaker rt‑MRI database and eight additional speakers show that the combined affine‑plus‑TPS approach with 12 and 14 landmarks yields the lowest mean point‑to‑closest‑point error of 3.19 mm.
By Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie
KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.
By Hyunjung Chung, Unsang Park
arXiv:2606.15848v2 Announce Type: replace
Abstract: 3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, an...
By Tingting Chen, Shaojun Wang, Huaye Zhang, Diqiong Jiang, Chenglizhao Chen
arXiv:2603.08249v2 Announce Type: replace-cross
Abstract: Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but...
By Pol Buitrago, Javier Hernando
arXiv:2609.22264v1 Announce Type: cross
Abstract: State-of-the-art models for audio-driven digital human generation have achieved photo-realistic results in talking-head synthesis. However, extending...
By Yichi Zhang, Hui Zhang, Guanjun Liu, Yuefeng Zou, Fengzhao Sun, Jun Yu
arXiv:2606. 16595v1 Announce Type: cross Abstract: Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible to language-specific variations.
By Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou
arXiv:2609.38019v1 Announce Type: new
Abstract: We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rat...
By Bangxun Tang
arXiv:2609.39273v1 Announce Type: new
Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...
By Junyao Gao, Sibo Liu, Weidong Zhang, Cairong Zhao, Jun Zhang