KoUniTalk is a lightweight, articulation‑centered benchmark that unifies Korean and English 3D talking‑face datasets onto a single mesh topology. By retargeting VOCASET and Korean speech‑based 3D data to a shared 1,176‑vertex template, it reduces output dimensionality from tens of thousands to 3,528 dimensions, focusing on the mouth and adjacent lower‑face regions. The benchmark includes 22 speakers, 4,978 sequences, and 642,781 frames, enabling controlled speech‑driven facial articulation training and cross‑dataset evaluation in a compact, identity‑neutral space.
By Hyunjung Chung, Unsang Park
The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.
By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing
Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.
By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2606. 26534v1 Announce Type: cross Abstract: Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.
By Tianxin Xie, Chenxing Li, Dong Yu, Li Liu
arXiv:2603. 04219v2 Announce Type: replace-cross Abstract: We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis.
By Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim