arXiv Computer Vision

The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation

The paper introduces a geometric measure of coarticulation for speech‑driven 3D facial animation, comparing lip‑path length to the shortest route through vowel, consonant, and vowel positions. Using only forced alignment, the measure evaluates four state‑of‑the‑art animation methods, revealing that all produce flatter lip trajectories than captured speech and that some methods lose 15–60% of the fast articulatory component. A pre‑registered viewer study confirms that damping real motion lowers perceived quality while exaggeration is not penalized, and viewers prefer real speech in 73.4% of sentence comparisons.

arXiv Machine Learning
Sep 29

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

The paper introduces a new framework for speech‑driven 3D facial animation that explicitly models visible articulatory dynamics. It uses a Speech‑Articulatory Memory (SAM) to link speech to three directional articulatory motions—spreading, opening, and protrusion—under phonetic context, and a Topology‑aware Articulatory Composition (TAC) to integrate these motions into surface‑consistent facial motion. Experiments on VOCASET and TFHP demonstrate state‑of‑the‑art reconstruction quality and improved lip articulation metrics, with a user study confirming better lip sync and realism.

By Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
arXiv Computer Vision
Sep 18

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

By Hyunjung Chung, Unsang Park
arXiv Computation and Language
Sep 24

Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond

The article surveys brain‑to‑language decoding, covering tasks, neural signals, methods, and evaluation across invasive and non‑invasive modalities. It traces the field’s evolution from constrained recognition to text generation, streaming speech, and facial animation, linking tasks to neural populations and decoder representations. The review highlights complementary decoding targets, shared representations, and the growing importance of calibration, feedback, and user control for online communication, while proposing a five‑level future trajectory toward bidirectional cognitive exchange.

By Yiqian Yang, Yiqun Duan, Chenyu Liu, Yiqi Wang, Xinliang Zhou, Chin-Teng Lin, Yu Zhang
arXiv Computer Vision
4d ago

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

The paper introduces Align Then Reason (ATR), a multilingual lip‑sync judge that aligns frame‑level lip representations with phonetic units of a candidate text line and then uses a language model to evaluate both content and timing. ATR achieves significant improvements over existing baselines on a seven‑language benchmark, with mean AUC gains of up to 59.4% for 2B reasoners and similar gains across other LLM families. The method also transfers well to unseen languages and outperforms lip‑reading baselines on real dubbing tasks such as dub‑line reranking and script‑to‑clip assignment.

By Rui Liu, Bhavin Jawade, Haoqi Li, Shivam Mehta, Karan Saxena, Yinghong Lan, Cameron R. Wolfe
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev