arXiv AI By Jiwon Kang, Heeji Yoon, Jaewoo Jung, Jaewon Min, Minkyeong Jeon, Biyeon Hwang, Sangwon Jung, Seungryong Kim

Transferability Between Understanding and Generation in Unified Multimodal Models

Read the original on arXiv AI →

arXiv:2607. 04423v1 Announce Type: cross Abstract: Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 5

Transferability Between Understanding and Generation in Unified Multimodal Models

Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision.

arXiv Computer Vision
Sep 2

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.

By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
arXiv AI
Sep 7

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study

The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.

By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng