arXiv Machine Learning By Luke Budny, Yuhong Guo, Kevin Cheung

Quality-Aware Modulation for Diffusion Transformers

Read the original on arXiv Machine Learning →

arXiv:2606. 30934v1 Announce Type: new Abstract: Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 8

DiffCVE: Diffusion-based Compressed Video Enhancement

Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Recent diffusion models have demonstrated strong generative capability for visual restoration, but directly applying them to compressed video often ignores compression degradation characteristics and may introduce structure-inconsistent hallucinations.