arXiv:2510.12041v3 Announce Type: replace
Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...
By Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang
arXiv:2507. 17853v2 Announce Type: replace-cross Abstract: Recent advances in text-to-image (T2I) generation have led to impressive visual results.
By Lifeng Chen, Jiner Wang, Zihao Pan, Beier Zhu, Xiaofeng Yang, Chi Zhang
The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.
By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
HyperErase introduces a hypernetwork-based framework for concept erasure in text-to-image models, replacing static adapters with prompt-conditioned parameter synthesis. The method maps textual descriptions to LoRA updates, eliminating per-prompt gradient optimization and manual merging. A decoupled rectification strategy further stabilizes and refines the synthesized adapters, yielding improved erasure effectiveness, image quality, and semantic alignment across diverse concepts.
By Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
arXiv:2606. 08492v1 Announce Type: cross Abstract: Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity of user prompts.
By Xuanyi Liu, Deyi Ji, Junyu Lu, Jing Wang, Qianxiong Xu, Xuhang Chen, Tianrun Chen, Siwei Ma
VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.
By Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu