The paper investigates image tokenizers as the visual language of unified multimodal models by creating a controlled autoregressive testbed that tracks task‑specific validation losses during multimodal continual pretraining across text, image, text‑to‑image, and image‑to‑text predictions. It shows that losses must be analyzed by task, that the loss–performance relationship varies with the token space, and that better reconstruction does not always lead to stronger downstream performance. The study also demonstrates how tokenizer design choices—such as discriminator use, semantic supervision, and vocabulary size—affect joint modeling and downstream results.
arXiv:2608.28609v1 Announce Type: cross
Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved...
By Bojie Li, Noah Shi
arXiv:2608. 05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning.
By Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu, Ponnurangam Kumaraguru
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
The study investigates whether multimodal large language models (MLLMs) report bistable images, like the duck‑rabbit, in a manner similar to humans. Using the LLaVA family, researchers examined two dimensions: modulability (the influence of visual cues and linguistic priors) and exclusivity (whether responses commit to a single interpretation). Results show that both visual and linguistic manipulations shift reports in human‑consistent ways while maintaining predominantly exclusive responses, driven by competing image‑token representations and distinct bottom‑up and top‑down pathways.
By Ryota Takatsuki, Tomoki Doi, Amane Watahiki, Anil K. Seth, Hitomi Yanaka
arXiv:2603. 04419v2 Announce Type: replace-cross Abstract: We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs).
By Murad Farzulla