arXiv AI

DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models

The paper introduces DART, a training‑free technique that combines low‑rank coordinate transport with target‑schedule response calibration to improve the reuse of LoRA adapters in few‑step video diffusion models. By using forward evaluations without source training videos, DART enhances joint quality scores and functional retention on a four‑step Wan2.2 target, with calibration contributing most of the gains. Component analysis shows complementary benefits from coordinate transport, and adapter‑level results indicate both positive functional effects and reduced negative transfer across different adapters.

Hugging Face Trending Papers
Sep 17

DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models

The paper introduces DART, a training‑free technique that combines low‑rank coordinate transport with target‑schedule response calibration to improve the reuse of LoRA adapters in few‑step video diffusion models. By avoiding source training videos and using forward evaluations, DART raises the joint quality score on a four‑step Wan2.2 target from 0.9029 to 0.9227 and shifts macro functional retention from negative to positive. Component analysis shows that calibration drives most of the quality gains, while coordinate transport adds complementary benefits, and the method demonstrates consistent improvements across additional targets.

arXiv AI
Sep 25

Accelerating Video Diffusion via Training-Free Trajectory Routing

The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.

By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh
arXiv Computer Vision
Sep 15

CrossDistill: Balancing Quality and Diversity via Trajectory-Level Hybrid Few-Step Distillation

CrossDistill is a trajectory-level hybrid few-step distillation framework for diffusion models that balances quality and diversity by splitting the sampling trajectory at a crossover point. The high-noise interval uses a trajectory-preserving objective to maintain global mode coverage, while the low-noise interval applies a distribution-matching objective to sharpen local details, with the two stages coupled through the crossover state. This noise-level scheduling policy, demonstrated on text-to-video and image-to-video diffusion models, expands the few-step quality-diversity frontier by preserving seed-level variation while achieving competitive visual fidelity.

By Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang
Hugging Face Trending Papers
Sep 24

Accelerating Video Diffusion via Training-Free Trajectory Routing

Accelerating Video Diffusion via Training-Free Trajectory Routing (TRACK) introduces a heterogeneous denoising strategy that switches between large and small diffusion models at selected steps, determined by a calibration process that measures disagreement between model predictions. By routing quality-sensitive steps to the large model and low-disagreement steps to the small model, TRACK achieves significant speedups—up to 2.73×—across several video diffusion benchmarks while maintaining comparable quality and diversity. The method requires no retraining, architectural changes, or online dual-model evaluation, making it a practical acceleration paradigm for video diffusion.

arXiv Computer Vision
4d ago

LongLive-Plug: Once-for-All Distillation for Video Generation

LongLive‑Plug is a once‑for‑all distillation framework that learns reusable LoRA adapters on a base video diffusion model, enabling training‑free, plug‑and‑play deployment to a wide range of downstream models. These adapters provide single‑pass classifier‑free guidance, few‑step sampling, and long‑context error correction for autoregressive generation, and remain effective even when downstream models add conditioning branches or expand output channels. The authors demonstrate that the approach works on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation.

By Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen