arXiv:2602.15382v3 Announce Type: replace-cross
Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...
By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.
arXiv:2609.01200v1 Announce Type: new
Abstract: When the visual encoder and the language decoder of a vision-language model (VLM) run on different compute nodes, the intermediate visual-token embeddi...
By Reza Heidari, Hamed R. Tavakoli, Juho Kannala
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei
NeuronEye is a plug‑in framework that builds a sparse, concept‑level neuron vocabulary from intermediate vision‑language model (VLM) representations and selectively activates query‑relevant visual concepts during inference. It decomposes vision‑token states into an overcomplete sparse basis organized by concept clusters, uses the language query to activate relevant clusters, localizes the corresponding image patches, and injects the focused evidence back into the vision tokens, while a suppression mechanism attenuates dominant perceptual directions. Experiments on Qwen2.5‑VL‑7B and LLaVA‑1.6‑7B show that NeuronEye improves CV‑Bench overall accuracy by +3.1, boosts Distance by +9.5, and raises BLINK Multi‑view by +8.3, indicating that sparse neuron vocabularies can act as active interfaces for concept‑level visual reasoning.
By Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao
arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.
By Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan
arXiv:2606. 14040v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are typically trained to reconstruct the \textbf{entire} residual stream through a sparse dictionary, implicitly assuming that all activation content is amenable to sparse, monosemantic decomposition.
By Ruixuan Deng, Zehao Jin, Zekun Wang, Zihan Dong
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney
The paper introduces FMSGOC, a framework that leverages visual‑linguistic foundation models to improve semantic and goal‑oriented communication for 6G. By transmitting a sparse set of semantic anchors and using a pretrained diffusion model for masked completion, it reduces overfitting and achieves high rate efficiency, reaching 0.039 BPP while maintaining strong semantic fidelity and robustness on unseen data.
By Boliang Liu, Wint Yi Poe, Riccardo Trivisonno, Giuseppe Caire
arXiv:2606. 09131v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) commonly inherit the deep, symmetric Transformer backbone designed for unimodal text modeling, and apply the same computation uniformly to image and language tokens.
By Siyuan Liu, Jinyang Wu
arXiv:2606.06158v2 Announce Type: replace
Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous...
By Kevin Dave, Sai Aditya Patkuri, Chhaya Kumar Das, Gouranga Bala, Rajeshkumar SA, R. Venkatesh Babu
arXiv:2609.38823v1 Announce Type: new
Abstract: Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly becau...
By Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen