FlashAR is a lightweight post‑training adaptation framework that converts a pre‑trained raster‑scan autoregressive image model into a highly parallel generator using two‑way next‑token prediction. It preserves the original training objective by keeping the horizontal head for row‑wise prediction and adding a lightweight vertical head for column‑wise prediction, with a learnable fusion gate to combine the two predictions. A two‑stage adaptation pipeline—first initializing the vertical head from the pre‑trained model and then jointly fine‑tuning—yields up to a 22.9× speedup for 512×512 image generation while using only 0.05% of the original training data.
By Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang
arXiv:2609.39222v1 Announce Type: new
Abstract: High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff be...
By Xu Huang, Ye Huang, Zijun Liao, Yuwei Niu, Xiaojie Li, Menghan Zhou, De Wen Soh, Xiaotong Li, Daquan Zhou
arXiv:2608.30782v1 Announce Type: new
Abstract: Real-world image super-resolution (Real-ISR) aims to preserve structures supported by the degraded observation while reconstructing perceptually realis...
By Bingtian Qiao, Yue Shi, Yong Guo, Wenjun Zhang, Jiezhang Cao
arXiv:2603.15132v3 Announce Type: replace
Abstract: While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the raw pixel m...
By Hainuo Wang, Mingjia Li, Xiaojie Guo
arXiv:2606. 16996v1 Announce Type: cross Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes.
By Tran Dinh Tien, Zhiqiang Shen
arXiv:2607. 27372v1 Announce Type: new Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages.
By Alexi Gladstone, Heng Ji, Yilun Du