The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.
By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
The paper introduces DARTS, a method for tuning decoder representations during model merging. It addresses representation bias in autoregressive decoders by using an entropy‑weighted L1 loss and a per‑position additive bias to correct errors that accumulate across token positions. Experiments on code generation, mathematical reasoning, and instruction following with Llama‑2‑7B show that DARTS improves performance over standard surgery while adding only 0.1% extra parameters.
By Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian
arXiv:2607. 13188v1 Announce Type: new Abstract: Human cognition does not separate understanding and generation.
By Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro V\'elez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
arXiv:2506. 01850v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs).
By Wayner Barrios, Andr\'es Villa, Juan Le\'on Alc\'azar, SouYoung Jin, Bernard Ghanem
arXiv:2605. 25820v2 Announce Type: replace Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel.
By Yulin Yuan, Hongshuo Zhao, Xiangming Meng
arXiv:2606.18974v3 Announce Type: replace
Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...
By Pengyu Li, Zhitao Gao, Lingling Zhang, Muye Huang, Yuanming Li, Fangzhi Xu, Jun Liu
arXiv:2604.16067v2 Announce Type: replace-cross
Abstract: Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continu...
By Guransh Singh
The paper introduces MATE, a reinforcement‑learning‑based post‑training framework for unified multimodal models that lets the generation and understanding branches challenge each other instead of cooperating. In MATE, each branch proposes candidate outputs that the other must reproduce, and the solver is trained on the worst‑handled candidate, creating an evolving adversarial loop without a separate adversary. Experiments on Janus‑Pro‑1B show that MATE improves generation and understanding metrics, including GenEval (+2.4), DPG‑Bench (+1.7), and an average of nine understanding benchmarks (+0.7), while enhancing consistency across image‑text cycles.
By Wentao Zhou, Weijie Gan, Jiayun Wang
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.
By Jack Boylan, Chris Hokamp
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other.
arXiv:2607. 06445v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarkable multimodal reasoning capabilities.
By Yoav Baron, Sara Dorfman, Roni Paiss, Daniel Cohen-Or, Or Patashnik