arXiv Computer Vision By Danzel Serrano, Przemyslaw Musialski

The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

Read the original on arXiv Computer Vision →

The paper introduces a geometric measure of coarticulation for speech‑driven 3D facial animation, comparing lip‑path length to the shortest route through vowel, consonant, and vowel positions. Using only forced alignment, the measure evaluates four state‑of‑the‑art animation methods, revealing that all produce flatter lip trajectories than captured speech and that some methods lose 15–60% of the fast articulatory component. A pre‑registered viewer study confirms that damping real motion lowers perceived quality while exaggeration is not penalized, and viewers prefer real speech in 73.4% of sentence comparisons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Sep 29

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.

By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv Computer Vision
Sep 18

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

By Hyunjung Chung, Unsang Park