arXiv AI By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

Read the original on arXiv AI →

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 2

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.

By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
arXiv Computer Vision
Sep 14

Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

ReVisIT is a train‑free framework that turns retrieved image‑label pairs into units of visual thought, combining structured class definitions, multimodal retrieval, and alternating user/assistant injection before joint decoding. On several benchmarks—including Fast Open MiniImageNet, Bongard‑OpenWorld, and the newly released MAAC‑Bench—ReVisIT achieves performance comparable to or surpassing large, trained models while using far fewer parameters. The approach demonstrates that high‑quality retrieval and a simple turns layer can provide a universal performance boost across diverse multimodal tasks.

By Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du