Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVS...
arXiv:2609.37831v1 Announce Type: new
Abstract: Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise gene...
By Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, Zhibo Chen
ScoutNeRV introduces a content‑adaptive initialization framework that speeds up the training of hierarchical grid‑based video implicit neural representations (INRs). By using a lightweight scout network to select a pre‑trained expert from a memory bank, the method transfers the expert’s grid and decoder parameters to a target HiNeRV model, achieving a 21.25 dB PSNR boost before fine‑tuning and reaching near‑baseline performance after only 37 epochs. The approach delivers a 9.25× wall‑clock speedup while maintaining competitive rate–distortion performance.
By Naser Alizada, Farhang Baghban, Hashem Pishkar, Ali Mousavi
arXiv:2510. 09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage.
By Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Yao Lu, Song Han
arXiv:2607. 14898v1 Announce Type: cross Abstract: Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video.
By Minguk Kang, Suha Kwak
The paper introduces a zero‑shot subject‑driven video generation framework that eliminates the need for per‑subject tuning and large subject‑video datasets. It achieves this by separating identity injection—learned from subject‑image pairs—and motion‑awareness preservation—maintained with a small set of arbitrary videos, and optimizes both with stochastic switching and dropout techniques. Using CogVideoX‑5B, the method adapts a single model with only 200K subject‑image pairs and 4,000 arbitrary videos in 288 A100 GPU hours, representing roughly 1% of the compute required by previous zero‑shot baselines while preserving subject fidelity and motion quality.
By Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo