arXiv Computer Vision By Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong

Form and Void: Entangled Composition through an Autonomous AI Agent

Read the original on arXiv Computer Vision →

The paper introduces FaV-A, a multimodal agent that generates positive and negative space compositions in a staged manner. It first creates a base object, analyzes its shape to identify negative‑space semantics, and then produces compositional instructions for the final image. Experiments show that this approach yields more visually coherent and semantically aligned compositions than direct zero‑shot multimodal large language model baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 16

Reasoning with Image Generation

The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.

By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
arXiv AI
Sep 28

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci