arXiv Computer Vision By Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

Read the original on arXiv Computer Vision →

The paper introduces the CO-AID dataset, which captures systematic defects in state‑of‑the‑art text‑to‑image models when prompts involve complex composition such as multiple entities and attributes. Researchers manually curated 651 reference images across people, hand, object, and scene categories, edited ChatGPT‑generated prompts to emphasize compositional factors, and generated AI images with three T2I models. A subjective study with 29 participants produced multi‑label defect annotations, enabling training of a deep model that predicts defects and improves image generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 9

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

arXiv:2606. 08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes.

By Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu
Hugging Face Trending Papers
Jun 4

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality.