arXiv AI

ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

arXiv:2606. 16327v1 Announce Type: cross Abstract: Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in scale.

arXiv Machine Learning
5d ago

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.

By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv AI
Sep 25

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

The paper introduces a geometric adaptation framework for cross‑speaker acoustic‑to‑articulatory inversion that leverages anatomical landmarks on vertebrae and dental structures. By applying an affine transformation followed by a thin‑plate spline deformation, the method maps predicted vocal‑tract contours from a fixed model to unseen speakers without retraining. Experiments on a single‑speaker rt‑MRI database and eight additional speakers show that the combined affine‑plus‑TPS approach with 12 and 14 landmarks yields the lowest mean point‑to‑closest‑point error of 3.19 mm.

By Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie
arXiv Computer Vision
Sep 18

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

By Hyunjung Chung, Unsang Park
arXiv Computer Vision
3d ago

MegaAvatar: Controllable Talking Avatar Generation

arXiv:2609.39273v1 Announce Type: new Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...

By Junyao Gao, Sibo Liu, Weidong Zhang, Cairong Zhao, Jun Zhang