Hugging Face Trending Papers

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet (VDN) introduces a hybrid attention mechanism for video diffusion models, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). The design updates memory once per frame, uses separate output projections and learnable gates to balance the two branches, and employs a staged teacher‑alignment recipe to integrate the new pathway into pretrained models. When applied to MiniMax H3, VDN achieves a 14.5× speedup, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs compared to the 50‑step dense baseline.

arXiv Machine Learning
Sep 18

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet (VDN) introduces a hybrid attention mechanism for livestream video generation, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). VDA updates memory once per frame, integrating spatial tokens, while separate output projections and learnable gates balance the two branches. Applied to MiniMax H3, VDN achieves a 14.5× speedup over the dense baseline, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs.

By Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng
arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv AI
Aug 28

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.

By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
arXiv Computer Vision
Sep 1

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

arXiv:2608.31106v1 Announce Type: new Abstract: Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We...

By Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
Hugging Face Trending Papers
Aug 27

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.