arXiv Machine Learning

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

arXiv:2606. 29095v1 Announce Type: cross Abstract: Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short clips and applied to long-horizon videos through chunked sliding-window inference, often causing temporal discontinuities at chunk boundaries.

arXiv Computer Vision
Aug 31

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.

By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
arXiv Computer Vision
Aug 28

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

Ring Forcing is an autoregressive video diffusion framework that enhances long‑term memory by enforcing retrieval from distant history through a ring‑structured training strategy. It introduces a compression and timestep composition method to extend effective historical span to minutes, and a sparse RoPE mechanism for scalable memory adaptation. Experiments show that Ring Forcing outperforms state‑of‑the‑art models in minutes‑long coherence and object permanence.

By Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan
arXiv AI
Aug 28

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.

By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
Hugging Face Trending Papers
Aug 27

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.