arXiv Machine Learning

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

arXiv:2608. 01298v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.

arXiv Computer Vision
6d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv Computer Vision
Sep 7

Importance-Aware Low-Rank Distillation of Diffusion Transformers

The paper introduces SVDtrunc, a two‑step block‑level compression method for Diffusion Transformers (DiTs) that allocates ranks across blocks, applies truncated SVD to the least important ones, and then fine‑tunes all blocks with modular knowledge distillation and a rectified‑flow objective. Experiments on FLUX.dev show that SVDtrunc achieves near‑full performance at 68% of the original parameters and remains competitive even at 57%, outperforming all competing approaches on GenEval, HPSv2, and DPG benchmarks. The method also works well without fine‑tuning, complementing step distillation and offering a practical path to efficient large‑scale generative models.

By Denis Zavadski, Sebastian Heid, Damjan Kal\v{s}an, Stefan Roth, Carsten Rother
arXiv AI
Jun 17

Rethinking Cross-Layer Information Routing in Diffusion Transformers

arXiv:2605. 20708v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited.

By Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou, Yunhe Li, Cuifeng Shen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang
arXiv AI
Jul 22

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

arXiv:2607. 19064v1 Announce Type: cross Abstract: Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy.

By Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu