Hugging Face Trending Papers

The Seriality Gap in Video Diffusion Models

Read the original on Hugging Face Trending Papers →

When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
2d ago

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation introduces a hybrid diffusion approach that first applies autoregressive diffusion at high noise levels and then switches to parallel diffusion at low noise levels. This method coordinates interdependent events to produce valid state transitions while reducing sampling time compared to fully serial generation. Implemented with a pixel‑space diffusion transformer and a LoRA‑fine‑tuned pretrained video model, S2PD outperforms bidirectional baselines in rule adherence and achieves greater temporal stability and sampling efficiency across games, physical simulations, and real video.

arXiv AI
Aug 28

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.

By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing