arXiv AI By Qirui Jiao, Daoyuan Chen, Yilun Huang, Xika Lin, Ying Shen, Yaliang Li

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

Read the original on arXiv AI →

arXiv:2505. 16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 3

Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling

The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.

By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai
arXiv Computer Vision
6d ago

HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models

HyperErase introduces a hypernetwork-based framework for concept erasure in text-to-image models, replacing static adapters with prompt-conditioned parameter synthesis. The method maps textual descriptions to LoRA updates, eliminating per-prompt gradient optimization and manual merging. A decoupled rectification strategy further stabilizes and refines the synthesized adapters, yielding improved erasure effectiveness, image quality, and semantic alignment across diverse concepts.

By Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
arXiv Computer Vision
2d ago

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.

By Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu