arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.
By Nikolai R\"ohrich, Isabell Hans, Felix Krause, Bj\"orn Ommer
arXiv:2606. 28406v1 Announce Type: new Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts.
By Davie Chen
arXiv:2608.27885v1 Announce Type: new
Abstract: Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not dire...
By Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon
arXiv:2512.17730v2 Announce Type: replace
Abstract: Detectors of AI-generated images tend to inherit the biases of the data they are trained on: models fitted to GAN imagery learn to treat GAN-specif...
By Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan, Duc-Tien Dang-Nguyen, Fakhri Karray
The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.
By Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han, Yushi Hu, Emily Dinan, Amita Kamath, Michal Drozdzal, Reyhane Askari-Hemmat, Luke Zettlemoyer, Marjan Ghazvininejad
Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While methods improve the FID and related metrics, it is increasingly unclear whether they reflect real progress in generative modeling.
arXiv:2602. 05951v2 Announce Type: replace-cross Abstract: Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation.
By Junwan Kim, Jiho Park, Seonghu Jeon, Seungryong Kim
PAPT++ is a risk‑aware adversarial generation‑training framework designed to improve single domain generalization. It learns diverse semantic reference images per class and uses them as denoising targets in classifier‑guided diffusion synthesis, thereby generating challenging yet semantically consistent samples. These samples are iteratively combined with source data to update the classifier, progressively exposing it to difficult variations and enhancing generalization performance on standard benchmarks.
By Zhipeng Xu, De Cheng, Xinyang Jiang, Lingfeng He, Huaijie Wang, Dongsheng Li, Nannan Wang, Xinbo Gao
arXiv:2609.30988v1 Announce Type: new
Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...
By Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha
arXiv:2609.37576v1 Announce Type: new
Abstract: With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to captur...
By Yu Zhao, Jiarui Wang, Huiyu Duan, Ye Zhao, Jutao Tang, Juntong Wang, Guangtao Zhai, Xiongkuo Min
arXiv:2511.19811v2 Announce Type: replace-cross
Abstract: Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive output...
By Debin Meng, Chen Jin, Zheng Gao, Yanran Li, Ioannis Patras, Georgios Tzimiropoulos
arXiv:2509. 24223v2 Announce Type: replace Abstract: Editing the content of an image with a pretrained text-to-image model remains challenging.
By Jianxin Zhang, Clayton Scott