arXiv Computer Vision

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

arXiv Machine Learning
5d ago

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.

By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv Computer Vision
Sep 7

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Motion-Omni is an end‑to‑end framework that jointly generates spoken dialogue and full‑body motion, producing speech, facial expressions, and hand, upper‑body, and lower‑body movements directly from the hidden states of a language model. The system requires joint training of the language model, speech generator, and motion generator to maintain audio‑motion alignment, and it is supervised using a scalable, model‑agnostic pipeline that pseudo‑labels 422,856 speech‑motion pairs. With a Qwen2.5‑7B‑Instruct backbone, Motion‑Omni‑Q7 achieves near‑cascade performance on motion metrics while being 5.4× faster, and it outperforms other non‑teacher cascades on beat correlation, diversity, and word error rate.

By Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo