arXiv:2609.10317v1 Announce Type: new
Abstract: Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directio...
By Yanru An, Ruiyan Wang, Wenwu Wei, Rui Bu, Qi Wang, Hongwei Hu, Zhengxue Cheng, Rong Xie, Li Song, Wenjun Zhang
Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.
By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended...
arXiv:2608.23383v1 Announce Type: new
Abstract: Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, foll...
By Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
The paper introduces Routed Forcing, a method that improves audio‑driven streaming avatar generation by selectively applying different distillation objectives to semantic regions and noise stages. It uses Data‑Forcing Distillation on person regions at high noise levels to restore motion diversity, while retaining Distribution Matching Distillation for mouth and background to keep lip sync and scene stability. Experiments show up to 45% better dynamics and 7–25% higher diversity compared to the previous Self Forcing approach.
By Zihan Su, Siwen Lu, Junhao Zhuang, Zeyue Xue, Haoyang Huang, Guanghao Li, Xiaofeng Tan, Chun Yuan, Nan Duan
arXiv:2609.36995v1 Announce Type: cross
Abstract: Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-relate...
By Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin, Yushi Huang, Xin Zhang, Yi Zhang, Yu Liu, Jun Zhang
TimeSteer introduces inference‑time speech scheduling for joint audio‑visual diffusion models, enabling users to place speech and visual articulation within specified time intervals without fine‑tuning the backbone. The method leverages two properties of the denoising process: a timing‑sensitive text‑to‑audio cross‑attention head that reveals each utterance’s source span, and a predicted clean latent that already organizes coupled speech and visual content. TimeSteer localizes each utterance’s source span and remaps the associated audio‑visual latent to the target interval, and the authors present SpeechShift as the first benchmark for interval‑level speech scheduling.
By Chao Zhou, Yiling Chen, Qi Chu, Tao Gong, Nenghai Yu, Tianyi We
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.
arXiv:2608. 08638v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools.
By Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang, Zhe Xu, Yiwei Zhao, Min Liang, Shuang Chen, Xipeng Qiu
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.
Encore is a new framework for generating long, synchronized audio‑video content. It splits the problem into local continuity, handled by iterative chunk‑wise synthesis with cross‑chunk context, and global consistency, enforced through reference audio‑video signals with shifted position embeddings. The Adaptive Signal Routing mechanism learns attention biases and residual scales to modulate the influence of each conditioning signal, enabling end‑to‑end joint audio‑video generation and infinite‑length inference.
By Shaohua Pan, Junbao Chen, Shengyi He, Jingfeng Xue, Wen Tao, Haocheng Feng, Siming Fan, Dongwei Pan, Yi Yang, Wei He, Hang Zhou