arXiv AI By Sherzod Hakimov, Mattia D'Agostini, Ivan Samodelkin, David Schlangen

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

Read the original on arXiv AI →

arXiv:2606. 01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns, making accumulated common ground directly observable as a rendered image.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 21

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.

By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo