arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.
By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi
arXiv:2606. 04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes.
By Zirui Wang, Junwei Yu, Adam Yala, David M. Chan, Joseph E. Gonzalez, Trevor Darrell
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
arXiv:2609.17790v1 Announce Type: new
Abstract: Pre-trained vision-language models (VLMs) exhibit strong cross-domain recognition performance even without additional training. However, this robustnes...
By Akanksha Singh, Vinod K. Kurmi
arXiv:2605. 16713v2 Announce Type: replace-cross Abstract: Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between.
By Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
arXiv:2606. 12217v1 Announce Type: cross Abstract: World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions.
By Lu Qiu, Yizhuo Li, Yi Chen, Yuying Ge, Yixiao Ge, Xihui Liu
D‑Scope is a framework that links the interpretation of sparse autoencoder (SAE) features in diffusion transformers (DiTs) to controllable image generation. It aggregates SigLIP‑2 embeddings of highly activating image patches into visual centroids, matches target text descriptions against these centroids, and retrieves individual features without per‑feature text annotations. The method provides visual evidence for each selection and uses spatially masked interventions to test decoder directions under fixed generation conditions, evaluated across 150 SAEs and a benchmark of 100 target concepts.
By Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2606. 16193v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret.
By Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang
arXiv:2606. 00275v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training.
By Zijie Zhou, Dandan Zhu, Hangxiangpan Wang, Heng Zhang, Huishen Jiao, Yi Zhao
The paper introduces Vision-of-Thought (VoT), a framework that inserts a discrete visual-thinking layer between vision‑language models (VLMs) and diffusion transformers (DiTs). Instead of using VLMs solely as text encoders, they act as multimodal planners that generate VoT tokens—high‑level visual plans such as objects and layouts—before pixel rendering. A specialized VoT tokenizer is trained with a closed‑loop objective combining VLM alignment, feature reconstruction, and vector‑quantization losses, ensuring the tokens are semantically readable by the VLM while preserving necessary visual information. Experimental results show that VoT improves semantic alignment and offers a structured, interpretable interface for controllable generation.
By Jingxiang Sun, Chao Liao, Zhengxiong Luo, Chaorui Deng, Chen-lin Zhang, Junke Wang, Ceyuan Yang, Haoqi Fan, Weilin Huang