arXiv Computer Vision

Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

arXiv Computer Vision
Sep 18

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.

By Hyunjung Chung, Unsang Park
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv Computer Vision
Sep 21

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.

By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei
arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma