Hugging Face Trending Papers

SwiftVR: Real-Time One-Step Generative Video Restoration

Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusion-based VR models remain difficult to deploy on consumer-grade GPUs due to two main bottlenecks: quadratic spatial attention at high resolutions and the latency-memory overhead of large video autoencoders.

arXiv Computer Vision
4d ago

FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.

By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.

arXiv AI
Sep 3

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

The paper introduces VoRTeC, a video compression framework that leverages a foundational flow model to encode latent video representations compactly and predict their positions along flow trajectories. By integrating multi‑scale priors and avoiding access to flow‑matching network parameters, VoRTeC achieves one‑step decoding with high perceptual fidelity, while maintaining temporal consistency through tail‑frame reuse and prior caching. Experiments show a 58% reduction in bit consumption compared to prior diffusion‑based methods and a decoding speed increase ranging from 3 to 197 times, reaching 13 FPS at 720p and 32 FPS at 480p.

By Yichong Xia, Qinhong Wu, Qinhong Wu, Jinpeng Wang, Zeyuan Chen, Haoqian Wang
arXiv Computer Vision
4d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
arXiv Machine Learning
Jul 23

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

arXiv:2607. 20125v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens.

By Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song
arXiv Computer Vision
4d ago

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

RelayVSR introduces a streaming video super‑resolution framework that combines a large generative model, which produces reference latents for sparse keyframes, with a lightweight Dual‑Memory Video Transformer that super‑resolves every frame using these references and low‑resolution input. The method employs Video‑Aware Reference Optimization (VARO), a reinforcement‑learning strategy that optimizes both system‑level video quality and reference‑level keyframe fidelity, outperforming direct joint training. On 1080p video, RelayVSR achieves 29.29 FPS with modest GPU memory usage, significantly faster and more efficient than the FlashVSR‑Tiny baseline.

By Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, Zhibo Chen