arXiv AI By Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu, Peijie Chen, Hongwu Ding, Xiong Zhang, Di Wu, Meng Meng, Jian Luan, Lin Li, Qingyang Hong

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

Read the original on arXiv AI →

SyncVoice is a new automatic video dubbing framework that adds a lightweight Text‑Visual Fusion Module to a pretrained TTS system, aligning visual features with linguistic representations to produce temporally synchronized speech. The approach avoids complex architectural changes and achieves state‑of‑the‑art performance on the LRS3 dataset in zero‑shot dubbing. When further trained on a large bilingual audio‑visual corpus, SyncVoice improves vocal fidelity while maintaining synchronization, enabling a single model to dub both Chinese and English videos.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 4

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.

By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu