arXiv Computer Vision
1d ago

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

RelayVSR introduces a streaming video super‑resolution framework that combines a large generative model, which produces reference latents for sparse keyframes, with a lightweight Dual‑Memory Video Transformer that super‑resolves every frame using these references and low‑resolution input. The method employs Video‑Aware Reference Optimization (VARO), a reinforcement‑learning strategy that optimizes both system‑level video quality and reference‑level keyframe fidelity, outperforming direct joint training. On 1080p video, RelayVSR achieves 29.29 FPS with modest GPU memory usage, significantly faster and more efficient than the FlashVSR‑Tiny baseline.

By Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, Zhibo Chen
arXiv Machine Learning
Sep 24

ScoutNeRV: Rapid Encoding of Grid-Based Video INRs via ScoutNet

ScoutNeRV introduces a content‑adaptive initialization framework that speeds up the training of hierarchical grid‑based video implicit neural representations (INRs). By using a lightweight scout network to select a pre‑trained expert from a memory bank, the method transfers the expert’s grid and decoder parameters to a target HiNeRV model, achieving a 21.25 dB PSNR boost before fine‑tuning and reaching near‑baseline performance after only 37 epochs. The approach delivers a 9.25× wall‑clock speedup while maintaining competitive rate–distortion performance.

By Naser Alizada, Farhang Baghban, Hashem Pishkar, Ali Mousavi
arXiv Computer Vision
Sep 3

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

The paper introduces a zero‑shot subject‑driven video generation framework that eliminates the need for per‑subject tuning and large subject‑video datasets. It achieves this by separating identity injection—learned from subject‑image pairs—and motion‑awareness preservation—maintained with a small set of arbitrary videos, and optimizes both with stochastic switching and dropout techniques. Using CogVideoX‑5B, the method adapts a single model with only 200K subject‑image pairs and 4,000 arbitrary videos in 288 A100 GPU hours, representing roughly 1% of the compute required by previous zero‑shot baselines while preserving subject fidelity and motion quality.

By Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo