arXiv:2608.28696v1 Announce Type: new
Abstract: Visual in-context learning (ICL) with multimodal large language models (MLLMs) is effective for fine-grained visual classification, but each retrieved...
By Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury
The study investigates how unified vision‑language models (VLMs) can simultaneously support visual understanding and generation. Using controlled benchmarks (SmartWatch and modified CelebA) that pair VQA, captioning, and text‑to‑image tasks, the authors evaluate several LLM‑based architectures built on SigLIP and VQ‑VAE visual spaces. Results show that mixed training can improve both understanding and generation, but the gains depend on how well the visual input and output spaces are aligned; misaligned or distorted visual spaces can weaken or reverse these benefits. The paper also demonstrates that balancing data across tasks and controlling attribute frequencies can help recover underrepresented visual concepts, and that the transfer is driven more by the base language model’s learned relationships than by visual adapters.
By Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng
arXiv:2606. 03871v1 Announce Type: cross Abstract: Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text.
By Luis Palacios, Lorenzo Basile, Diego Doimo, Alberto Cazzaniga
arXiv:2511. 16107v3 Announce Type: replace-cross Abstract: Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training.
By Shao-Jun Xia, Huixin Zhang, Zhengzhong Tu
The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.
By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
arXiv:2608.22174v1 Announce Type: new
Abstract: Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding...
By Yubo Zhu, Zhehan Kan, Jingyi Yang, Miaolin Chen, Jinbo Xing, Kai Zhu, Zijian Wang, Sheng Zhong, Wei Tong
Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference.
The paper demonstrates that vision‑language models (VLMs) possess a small set of attention heads, called Visual Retrieval Heads (VRHs), that are causally responsible for linking text prompts to specific image regions. By adapting head‑scoring techniques from language models, the authors identify VRHs as the heads whose attention from output prediction tokens, summed over the ground‑truth referent region, most reliably indicates causal grounding. Experiments across eleven VLMs and five referring‑expression benchmarks show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect, and that VRHs generalize across diverse visual tasks and transfer across models sharing an LLM backbone.
arXiv:2606. 28406v1 Announce Type: new Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts.
By Davie Chen
arXiv:2603. 01195v2 Announce Type: replace-cross Abstract: The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning.
By Mingkang Dong, Hongyi Cai, Jie Li, Sifan Zhou, Bin Ren, Kunyu Peng, Yuqian Fu
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.
By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu