Hugging Face Trending Papers

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas.

arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv Computer Vision
Aug 25

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

arXiv:2608.23383v1 Announce Type: new Abstract: Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, foll...

By Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
arXiv Computer Vision
6d ago

Where and When to Force: Routed Forcing for Streaming Avatars

The paper introduces Routed Forcing, a method that improves audio‑driven streaming avatar generation by selectively applying different distillation objectives to semantic regions and noise stages. It uses Data‑Forcing Distillation on person regions at high noise levels to restore motion diversity, while retaining Distribution Matching Distillation for mouth and background to keep lip sync and scene stability. Experiments show up to 45% better dynamics and 7–25% higher diversity compared to the previous Self Forcing approach.

By Zihan Su, Siwen Lu, Junhao Zhuang, Zeyue Xue, Haoyang Huang, Guanghao Li, Xiaofeng Tan, Chun Yuan, Nan Duan
arXiv AI
Sep 2

TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models

TimeSteer introduces inference‑time speech scheduling for joint audio‑visual diffusion models, enabling users to place speech and visual articulation within specified time intervals without fine‑tuning the backbone. The method leverages two properties of the denoising process: a timing‑sensitive text‑to‑audio cross‑attention head that reveals each utterance’s source span, and a predicted clean latent that already organizes coupled speech and visual content. TimeSteer localizes each utterance’s source span and remaps the associated audio‑visual latent to the target interval, and the authors present SpeechShift as the first benchmark for interval‑level speech scheduling.

By Chao Zhou, Yiling Chen, Qi Chu, Tao Gong, Nenghai Yu, Tianyi We
Hugging Face Trending Papers
Aug 6

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.

Hugging Face Trending Papers
Aug 27

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.

arXiv Computer Vision
Sep 7

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Encore is a new framework for generating long, synchronized audio‑video content. It splits the problem into local continuity, handled by iterative chunk‑wise synthesis with cross‑chunk context, and global consistency, enforced through reference audio‑video signals with shifted position embeddings. The Adaptive Signal Routing mechanism learns attention biases and residual scales to modulate the influence of each conditioning signal, enabling end‑to‑end joint audio‑video generation and infinite‑length inference.

By Shaohua Pan, Junbao Chen, Shengyi He, Jingfeng Xue, Wen Tao, Haocheng Feng, Siming Fan, Dongwei Pan, Yi Yang, Wei He, Hang Zhou