Hugging Face Trending Papers
Sep 8

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

The paper investigates image tokenizers as the visual language of unified multimodal models by creating a controlled autoregressive testbed that tracks task‑specific validation losses during multimodal continual pretraining across text, image, text‑to‑image, and image‑to‑text predictions. It shows that losses must be analyzed by task, that the loss–performance relationship varies with the token space, and that better reconstruction does not always lead to stronger downstream performance. The study also demonstrates how tokenizer design choices—such as discriminator use, semantic supervision, and vocabulary size—affect joint modeling and downstream results.

arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv AI
3d ago

(How) Do MLLMs Report Bistable Images Like Humans?

The study investigates whether multimodal large language models (MLLMs) report bistable images, like the duck‑rabbit, in a manner similar to humans. Using the LLaVA family, researchers examined two dimensions: modulability (the influence of visual cues and linguistic priors) and exclusivity (whether responses commit to a single interpretation). Results show that both visual and linguistic manipulations shift reports in human‑consistent ways while maintaining predominantly exclusive responses, driven by competing image‑token representations and distinct bottom‑up and top‑down pathways.

By Ryota Takatsuki, Tomoki Doi, Amane Watahiki, Anil K. Seth, Hitomi Yanaka