Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
arXiv:2602. 12304v5 Announce Type: replace-cross Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts.
By Maomao Li, Zhen Li, Kaipeng Zhang, Guosheng Yin, Zhifeng Li, Dong Xu
Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas.
arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.
By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations.
arXiv:2606. 01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos.
By Zhicheng Zhang, Lei Wang, Yu Zhang, Yongsheng Gao
arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.
By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.
By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
arXiv:2609.10394v1 Announce Type: cross
Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...
By Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv:2608.31106v1 Announce Type: new
Abstract: Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We...
By Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.
By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei